EDBT 2026 Demo / reviewers in the wild / expert
Jürgen Beyerer
dblp:00/5069 · also Juergen Beyerer
· DBLP profile ↗
122ranked-venue papers
2as first author
48since 2021 · last 2026
0000-0003-3556-7181ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 67 · 2 first-author · 22 since 2021Artificial intelligence and machine learning · 40 · 26 since 2021Systems, architecture and hardware · 12 · 6 since 2021Security and privacy · 12 · 6 since 2021Human-computer interaction and ubiquitous computing · 9Databases, data management, data science and information retrieval · 8 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uplifting 2D to 3D Human Poses with Joint Rotations and Bone Constraints: A Strong Baseline for Sports and Fitness ApplicationsabstractAccurate 3D human pose estimation plays a pivotal role in AI-driven systems for assessing and improving skilled human activities in domains such as fitness and sports. Despite significant progress in 2D-to-3D pose uplifting, current methods often focus on isolated components, such as kinematics modeling, bone-length constraints, or joint orientation estimation. Surprisingly, little effort has been made to combine these complementary approaches into a unified framework, even though their integration could naturally enhance performance. To address this, we introduce a strong baseline for 2D-to-3D pose uplifting that combines quaternion-based joint orientation estimation, bone length constraints for anatomically plausible predictions, and kinematic modeling for improved stability. Our approach is evaluated on the traditional Human3.6M dataset, and three challenging, underexplored benchmarks: Fit3D, which focuses on fitness exercises, AthletePose3D, designed to study professional athletic performance, and Harmony4D with complex inter-person interactions. These datasets pose unique challenges, such as high-speed and non-regular motions, which are more representative of real-world applications than conventional datasets. Results demonstrate that integrating these complementary techniques improves 3D pose estimation accuracy and provides a robust foundation for downstream tasks such as skill assessment and feedback generation. By addressing the challenges posed by complex datasets and highlighting the advantages of combining well-established methods, this work aims to inspire further research in skilled activity understanding. Our strong baseline provides a lightweight and effective baseline for 3D human pose estimation in fitness and sports contexts. The code is available at: https://github. com/Mickael-Cormier/3d-hpe-strong-baseline. Mickael Cormier, Jeremy Zolk, Astrid Laubenheimer, Jürgen Beyerer |
FG | 4 |
| 2026 | Higher-Order Adversarial Patches for Real-Time Object Detectors
Jens Bayer, Stefan Becker, David Münch, Michael Arens, Jürgen Beyerer |
ICPR (2) | 5 |
| 2026 | Rethinking Hierarchical Supervision: Revisiting Simplicity in the Era of Strong Visual Backbones
Philipp Thelen, Jürgen Beyerer |
ICPR (3) | 3 |
| 2026 | Scalable Video Action Anticipation with Cross Linear Attentive MemoryabstractRecent advances in action anticipation rely heavily on Transformer architectures to learn discriminative representations of the past observation, incurring high computational and memory overhead that limits their applicability to long videos. While temporal processors with linear complexity like RNNs and state-space models offer efficient alternatives, their sequential nature risks overlooking subtle cues in observed frames that could enhance future anticipation. We address this limitation with Cross Linear Attentive Memory (CLAM), a memory module that selectively retrieves complementary context cues from frame features. By reformulating linear attention to replace traditional cross-attention, CLAM achieves linear computation complexity and constant memory usage relative to input length. Finally, by fusing the outputs of the temporal processor and CLAM, a non-autoregressive Transformer decoder generates future actions in one shot with high accuracy. Experiments on egocentric (EpicKitchens100 and Ego4D) and third-person (Thumos14) benchmarks demonstrate our model’s superior anticipation accuracy and scalability, processing longer sequences with significantly less latency growth than alternatives. Our approach also achieves promising results in online action detection. Zeyun Zhong, Manuel Martin, David Schneider 0006, David J. Lerch, Chengzhi Wu, Frederik Diederichs, Juergen Gall, Jürgen Beyerer |
WACV | 8 |
| 2025 | SAMBLE: Shape-Specific Point Cloud Sampling for an Optimal Trade-Off Between Local Detail and Global UniformityabstractDriven by the increasing demand for accurate and efficient representation of 3D data in various domains, point cloud sampling has emerged as a pivotal research topic in 3D computer vision. Recently, learning-to-sample methods have garnered growing interest from the community, particularly for their ability to be jointly trained with downstream tasks. However, previous learning-based sampling methods either lead to unrecognizable sampling patterns by generating a new point cloud or biased sampled results by focusing excessively on sharp edge details. Moreover, they all overlook the natural variations in point distribution across different shapes, applying a similar sampling strategy to all point clouds. In this paper, we propose a Sparse Attention Map and Bin-based Learning method (termed SAMBLE) to learn shape-specific sampling strategies for point cloud shapes. SAMBLE effectively achieves an improved balance between sampling edge points for local details and preserving uniformity in the global shape, resulting in superior performance across multiple common point cloud downstream tasks, even in scenarios with few-point sampling. Chengzhi Wu, Yuxin Wan, Julius Pfrommer, Zeyun Zhong, Junwei Zheng, Jürgen Beyerer |
CVPR | 8 |
| 2025 | 3D Extended Object Tracking Based on Extruded B-Spline Side View ProfilesabstractObject tracking is an essential task for autonomous systems. With the advancement of 3D sensors, these systems can better perceive their surroundings using effective 3D Extended Object Tracking (EOT) methods. Based on the observation that common road users are symmetrical on the right and left sides in the traveling direction, we focus on the side view profile of the object. In order to leverage of the development in 2D EOT and balance the number of parameters of a shape model in the tracking algorithms, we propose a method for 3D extended object tracking (EOT) by describing the side view profile of the object with B-spline curves and forming an extrusion to obtain a 3D extent. The use of B-spline curves exploits their flexible representation power by allowing the control points to move freely. The algorithm is developed into an Extended Kalman Filter (EKF). For a through evaluation of this method, we use simulated traffic scenario of different vehicle models and real-world open dataset containing both radar and lidar data. Longfei Han, Klaus Kefferpütz, Jürgen Beyerer |
FUSION | 3 |
| 2025 | How Prompting Shapes Decisions: Analyzing LLM Behavior in XAI-Augmented Decision Support Systems
Finn Schwall, Maximilian Becker, Anmol Ashri, Jürgen Beyerer |
IJCCI (3) | 4 |
| 2025 | Comparison of Vehicle Lateral Movement Models for Automated Driving Function ValidationabstractThe range of vision of vehicle sensors used by automated driving functions is considerably influenced by the lateral movement of vehicles within their lane. With the increasing relevance of simulations for the validation of automated driving functions, realistic modeling of this lateral movement gains in importance. Different stochastic models addressing this task have been proposed in literature. The used datasets and performed evaluations, however, are diverse, hampering the comparison of the models. Further, in some cases the evaluated features are limited, restricting the assessment of the model's applicability for a wide range of use cases. This work therefore starts with the identification of a suitable uniform data basis and evaluation strategy enabling the latter. Based on this, the models are compared in terms of qualitative and quantitative results and runtime. The results reveal strengths and limitations of the models, indicate suitable use cases, and reveal potentials for improvement. Nicole Neis, Jens R. Ziehn, Masoud Roschani, Jürgen Beyerer |
IV | 4 |
| 2025 | Towards Graph-based Self-learning of Industrial Process Behaviour for Anomaly DetectionabstractThe increasing sophistication of cyber threats targeting industrial control systems (ICS) necessitates advanced anomaly detection techniques capable of identifying attacks by analyzing industrial process data exchange. This paper addresses the challenge of representing and learning the spatio-temporal characteristics of industrial network communication as Graph for self-learning anomaly detection. We propose a novel framework that models the spatio-temporal characteristics as graph snapshots and applies Graph Neural Networks (GNNs) for anomaly detection. Each graph snapshot captures the structural and temporal dynamics of PROFINET-based industrial traffic, with edge features encoding the payload transitions, timing intervals, and cycle counter differences. We evaluate the performance of an isotropic Graph Convolutional Network (GCN) and anisotropic GNN variants – Message Passing Neural Network (MPNN), Gated Graph ConvNet (GatedGCN) and Graph Transformer (GT) – for the task of graph classification on real-world datasets from a miniaturized deterministic production plant. Our evaluation results demonstrate that anisotropic GNN models (MPNN, GatedGCN, GT) achieve complete anomaly detection (specificity) while maintaining perfect recall on normal behavior (sensitivity). In contrast, the isotropic GCN fails to distinguish between normal and anomalous states of the miniaturized plant. These findings highlight the efficacy of encoding spatio-temporal characteristics on graph edges and the capability of anisotropic GNNs to learn complex process behaviors for anomaly detection in the industrial networks of the evaluated production plant. Ankush Meshram, Markus Karch, Christian Haas 0005, Jürgen Beyerer |
KES | 4 |
| 2025 | Traversing the subspace of adversarial patchesabstractAbstract Despite ongoing research on the topic of adversarial examples in deep learning for computer vision, some fundamentals of the nature of these attacks remain unclear. As the manifold hypothesis posits, high-dimensional data tends to be part of a low-dimensional manifold. To verify the thesis with adversarial patches–a special form of adversarial attack that can be used to fool object detectors in the physical world–this paper provides an analysis of a set of adversarial patches and investigates the reconstruction abilities of five different dimensionality reduction methods. Quantitatively, the performance of reconstructed patches in an attack setting is measured and the impact of sampled patches from the latent space during adversarial training is investigated. The evaluation is performed on two publicly available datasets for person detection. The results indicate that more sophisticated dimensionality reduction methods offer no advantages over a simple principal component analysis. Jens Bayer, Stefan Becker, David Münch, Michael Arens, Jürgen Beyerer |
Mach. Vis. Appl. | 5 |
| 2024 | A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised LearningabstractContrastive learning is an essential method in self-supervised learning. It primarily employs a multi-branch strategy to compare latent representations obtained from different branches and train the encoder. In the case of multi-modal input, diverse modalities of the same object are fed into distinct branches. When using single-modal data, the same input undergoes various augmentations before being fed into different branches. However, all existing contrastive learning frameworks have so far only performed contrastive operations on the learned features at the final loss end, with no information exchange between different branches prior to this stage. In this paper, for point cloud unsupervised learning without the use of extra training data, we propose a Contrastive Cross-branch Attention-based framework for Point cloud data (termed PoCCA), to learn rich $3 D$ point cloud representations. By introducing sub-branches, PoCCA allows information exchange between different branches before the loss end. Experimental results demonstrate that in the case of using no extra training data, the representations learned with our self-supervised model achieve state-of-the-art performances when used for downstream tasks on point clouds. Chengzhi Wu, Qianliang Huang, Julius Pfrommer, Jürgen Beyerer |
3DV | 5 |
| 2024 | DDS Security+: Enhancing the Data Distribution Service With TPM-based Remote AttestationabstractThe Data Distribution Service (DDS) is a widely accepted industry standard for reliably exchanging data over the network using a publish-subscribe model. While DDS already includes basic security features such as participant authentication and access control, the possibilities of leveraging Trusted Platform Modules (TPMs) to increase the security and trustworthiness of DDS-based applications have not been sufficiently researched yet. In this work, we show how TPM-based remote attestation can be effectively integrated into the existing DDS security architecture. This enables application developers to verify the code integrity of remote DDS participants during the operation of the distributed system. Our solution transparently extends the DDS secure channel handshake, while cryptographically binding the established communication channels to the attested software stacks. We show the security properties of our proposal by formally verifying the resulting remote attestation protocol using the Tamarin theorem prover. We also implement our solution as a fork of the popular eProsima FastDDS library and evaluate the resulting performance impact when conducting TPM-based remote attestations of DDS applications. Paul Georg Wagner, Pascal Birnstill, Jürgen Beyerer |
ARES | 3 |
| 2024 | Rethinking Attention Module Design for Point Cloud Analysis
Chengzhi Wu, Kaige Wang, Zeyun Zhong, Junwei Zheng, Julius Pfrommer, Jürgen Beyerer |
ICPR (26) | 8 |
| 2024 | SynthAct: Towards Generalizable Human Action Recognition based on Synthetic DataabstractSynthetic data generation is a proven method for augmenting training sets without the need for extensive setups, yet its application in human activity recognition is underexplored. This is particularly crucial for human-robot collaboration in household settings, where data collection is often privacy-sensitive. In this paper, we introduce SynthAct, a synthetic data generation pipeline designed to significantly minimize the reliance on real-world data. Leveraging modern 3D pose estimation techniques, SynthAct can be applied to arbitrary 2D or 3D video action recordings, making it applicable for uncontrolled in-the-field recordings by robotic agents or smarthome monitoring systems. We present two SynthAct datasets: AMARV, a large synthetic collection with over 800k multi-view action clips, and Synthetic Smarthome, mirroring the Toyota Smarthome dataset. SynthAct generates a rich set of data, including RGB videos and depth maps from four synchronized views, 3D body poses, normal maps, segmentation masks and bounding boxes. We validate the efficacy of our datasets through extensive synthetic-to-real experiments on NTU RGB+D and Toyota Smarthome. SynthAct is available on our project page4. David Schneider 0006, Marco Keller, Zeyun Zhong, Kunyu Peng, Alina Roitberg, Jürgen Beyerer, Rainer Stiefelhagen |
ICRA | 6 |
| 2024 | Synset Boulevard: A Synthetic Image Dataset for VMMR*abstractWe present and discuss the Synset Boulevard dataset, designed for the task of surveillance-nature vehicle make and model recognition (VMMR)—to the best of our knowledge the first entirely synthetically generated large-scale VMMR image dataset. Through the simulation of image data rather than the manual annotation of real data, we intend to mitigate common challenges in state-of-the-art VMMR datasets, namely bias, human error, privacy, and the challenge of providing systematic updates. On the other hand, the provision and use of synthetic data introduce individual challenges, such as potential domain gaps and a less pronounced intra-class variance. Our approach to address these challenges, using path tracing and physically-based, data-driven models, is evaluated on an existing large real-world dataset. Overall, our synthetic dataset contains 32 400 independent images (each with different imaging simulations and with/without masked license plates, leading to a total of 259 200 images) from 162 different vehicle models of 43 makes depicted in front view. It is split into 8 sub-datasets to investigate the influence of optical/imaging effects on the classification ability. Anne Sielemann, Masoud Roschani, Jens R. Ziehn, Jürgen Beyerer |
ICRA | 5 |
| 2024 | Joint Parameter and State-Space Modelling of Manufacturing Processes using Gaussian ProcessesabstractManufacturing process optimization is an open question, where Bayesian decision theoretic methods have shown considerable promise. One such is Bayesian optimization, with Gaussian Process (GP) surrogate model. This paper explores Gaussian Processes networks to jointly use parameter and observed state to predict the output(s) of a manufacturing process. The Gaussian process network that represents the paths from parameters to state-space to tasks, provides a methodology to ‘look inside’ the black-box of complex manufacturing processes. We present a comparative analysis of this method against the multi-task Gaussian processes and single-task counterparts, highlighting the benefits and drawbacks of each in modelling the behavior of such processes. We show the benefits of the proposed approach using numerical experiments. We show that we are able to improve the output prediction by additional sensor observations from inside the process at training time without needing those sensor observations for predicting product quality given the process parameters. Saksham Kiroriwal, Julius Pfrommer, Hendrik Mende, Robert H. Schmitt, Jürgen Beyerer |
INDIN | 5 |
| 2024 | Scalable Radar-based Roadside Perception: Self-localization and Occupancy Heat Map for Traffic Analysisabstract4D mmWave radar sensors are suitable for roadside perception in city-scale Intelligent Transportation Systems (ITS) due to their long sensing range, weatherproof functionality, simple mechanical design, and low manufacturing cost. In this work, we investigate radar-based ITS for scalable traffic analysis. Localization of these radar sensors at city scale is a fundamental task in ITS. For flexible sensor setups, it requires even more effort. To address this task, we propose a self-localization approach that matches two descriptions of the "road": the one from the geometry of the motion trajectories of cumulatively observed vehicles, and the other one from the aerial laser scan. An Iterative Closest Point (ICP) algorithm is used to register the motion trajectory in the road section of the laser scan. The resulting estimate of the transformation matrix represents the sensor pose in a global reference frame. We evaluate the results and show that the method outperforms other map-based radar localization methods, especially for the orientation estimation. Beyond the localization result, we project radar sensor data onto a city-scale laser scan and generate a scalable occupancy heat map as a traffic analysis tool. This is demonstrated using two radar sensors monitoring an urban area in the real world. Longfei Han, Qiuyu Xu, Klaus Kefferpütz, Gordon Elger, Jürgen Beyerer |
IV | 6 |
| 2024 | Few-Shot Semantic Segmentation for Complex Driving ScenesabstractThe main objective of few-shot semantic segmentation (FSSS) is to segment novel objects within query images by leveraging a limited set of support images. Being capable of segmenting the novel classes plays an essential role in the development of perception functions for automated vehicles. However, existing few-shot semantic segmentation work strives to improve the performance of the models on object-centric datasets. In our work, we evaluate the few-shot semantic segmentation on the more challenging driving scene understanding tasks. As a use case specific study, we give a systematic analysis of the disparity between commonly used FSSS datasets and driving datasets. Based on that, we proposed methodologies to integrate knowledge from the class hierarchy of the datasets, utilize more effective feature extraction, and choose more representative support images during inference. These approaches are evaluated extensively on the Cityscapes and Mapillary datasets to indicate their effectiveness. We point out the remaining challenges of training, evaluating, and employing FSSS models for complex road scenes in real practice. Jingxing Zhou, Ruei-Bo Chen, Jürgen Beyerer |
IV | 3 |
| 2024 | Self-Supervised Generative-Contrastive Learning of Multi-Modal Euclidean Input for 3D Shape Latent Representations: A Dynamic Switching ApproachabstractWe propose a combined generative and contrastive neural architecture for learning latent representations of 3D volumetric shapes. The architecture uses two encoder branches for voxel grids and multi-view images from the same underlying shape. The main idea is to combine a contrastive loss between the resulting latent representations with an additional reconstruction loss. That helps to avoid collapsing the latent representations as a trivial solution for minimizing the contrastive loss. A novel dynamic switching approach is used to cross-train two encoders with a shared decoder. The switching approach also enables the stop gradient operation on a random branch. Further classification experiments show that the latent representations learned with our self-supervised method integrate more useful information from the additional input data implicitly, thus leading to better reconstruction and classification performance. Chengzhi Wu, Julius Pfrommer, Mingyuan Zhou, Jürgen Beyerer |
IEEE Trans. Multim. | 4 |
| 2023 | NIFF: Alleviating Forgetting in Generalized Few-Shot Object Detection via Neural Instance Feature ForgingabstractPrivacy and memory are two recurring themes in a broad conversation about the societal impact of AI. These con-cerns arise from the need for huge amounts of data to train deep neural networks. A promise of Generalized Few-shot Object Detection (G-FSOD), a learning paradigm in AI, is to alleviate the need for collecting abundant training samples of novel classes we wish to detect by leveraging prior knowledge from old classes (i.e., base classes). G-FSOD strives to learn these novel classes while alleviating catas-trophic forgetting of the base classes. However, existing approaches assume that the base images are accessible, an assumption that does not hold when sharing and storing data is problematic. In this work, we propose the first data-free knowledge distillation (DFKD) approach for G-FSOD that leverages the statistics of the region of interest (RoI) features from the base model to forge instance-level features without accessing the base images. Our contribution is three-fold: (1) we design a standalone lightweight generator with (2) class-wise heads (3) to generate and replay diverse instance-level base features to the RoI head while finetuning on the novel data. This stands in contrast to standard DFKD approaches in image classification, which invert the entire network to generate base images. Moreover, we make careful design choices in the novel finetuning pipeline to regularize the model. We show that our approach can dramatically reduce the base memory requirements, all while setting a new standard for G-FSOD on the challenging MS-COCO and PASCAL-VOC benchmarks. Karim Guirguis, Johannes Meier, George Eskandar, Matthias Kayser, Bin Yang 0009, Jürgen Beyerer |
CVPR | 6 |
| 2023 | Principles of Forgetting in Domain-Incremental Semantic Segmentation in Adverse Weather ConditionsabstractDeep neural networks for scene perception in automated vehicles achieve excellent results for the domains they were trained on. However, in real-world conditions, the domain of operation and its underlying data distribution are subject to change. Adverse weather conditions, in particular, can significantly decrease model performance when such data are not available during training. Additionally, when a model is incrementally adapted to a new domain, it suffers from catastrophic forgetting, causing a significant drop in performance on previously observed domains. Despite recent progress in reducing catastrophic forgetting, its causes and effects remain obscure. Therefore, we study how the representations of semantic segmentation models are affected during domain-incremental learning in adverse weather conditions. Our experiments and representational analyses indicate that catastrophic forgetting is primarily caused by changes to low-level features in domain-incremental learning and that learning more general features on the source domain using pre-training and image augmentations leads to efficient feature reuse in subsequent tasks, which drastically reduces catastrophic forgetting. These findings highlight the importance of methods that facilitate generalized features for effective continual learning algorithms. Tobias Kalb, Jürgen Beyerer |
CVPR | 2 |
| 2023 | Attention-Based Point Cloud Edge SamplingabstractPoint cloud sampling is a less explored research topic for this data representation. The most commonly used sampling methods are still classical random sampling and farthest point sampling. With the development of neural networks, various methods have been proposed to sample point clouds in a task-based learning manner. However, these methods are mostly generative-based, rather than selecting points directly using mathematical statistics. Inspired by the Canny edge detection algorithm for images and with the help of the attention mechanism, this paper proposes a non-generative Attention-based Point cloud Edge Sampling method (APES), which captures salient points in the point cloud outline. Both qualitative and quantitative experimental results show the superior performance of our sampling method on common benchmark tasks. Chengzhi Wu, Junwei Zheng, Julius Pfrommer, Jürgen Beyerer |
CVPR | 4 |
| 2023 | Towards Self-learning Industrial Process Behaviour from Payload Bytes for Anomaly DetectionabstractNetwork Intrusion Detection System (NIDS) for process-based anomaly detection have been developed as one of the cybersecurity solutions against industrial process targeted attacks such as Stuxnet. In practice, the real-world industrial plants could not complement the advancements in the industrial cybersecurity research as upgrading the infrastructure is an expensive and deterrent process for plant owners. In addition, the infrastructure information might be lost over the intended longer lifetime, hence, configuring a NIDS in the absence of such information is a challenge. Moreover, the existing NIDS solutions analyze the industrial process values/parameters with the knowledge of their semantics, and would fail when the semantics is not known or lost. As a solution to aforementioned problem, we propose an industrial communication paradigm aware Process Payload Profiling Framework (P3F), capable of self-learning process behavior from network traffic without the knowledge of underlying process parameters being exchanged. We also report P3F’s successful detection of an anomaly in the process of a miniaturized PROFINET-based industrial system, caused by a simulated process-targeted cyberattack. Ankush Meshram, Markus Karch, Christian Haas 0005, Jürgen Beyerer |
ETFA | 4 |
| 2023 | Past Information Aggregation for Multi-Person TrackingabstractMulti-person tracking is often solved with the tracking-by-detection (TBD) paradigm. So-far tracked targets are matched to new detections on the basis of the current track states including motion or appearance information. If tracks are updated with inaccurate detections or unreliable appearance features, e.g., due to occlusion, the track states can become distorted leading to errors in the association. To mitigate this problem, we propose a simple and generic method for Past Information Aggregation (PIA) that utilizes more tracking information from the past in the association and can be applied within any TBD approach. We combine PIA with a new sophisticated distance measure fusing motion and appearance cues and a second matching stage which further improves the association accuracy. Our tracking framework is analyzed with extensive ablative experiments and state-of-the-art results are achieved on the MOT17 and MOT20 benchmarks. Daniel Stadler, Jürgen Beyerer |
ICIP | 2 |
| 2023 | SWaTEval: An Evaluation Framework for Stateful Web Application Testingabstract430 Anne Borcherding, Nikolay Penkov, Mark Giraud, Jürgen Beyerer |
ICISSP | 4 |
| 2023 | Counterfactual Root Cause Analysis via Anomaly Detection and Causal GraphsabstractAnomalies in production processes can cause expensive standstills, damages to the production equipment, waste of materials and flaws in the final product. In production, finding anomalies is usually accomplished by machine learning methods. But to avert anomalies and to automatically recover, actually the detection of the root causes is required. We developed an approach that detects anomalies and then deduces root causes by combining an anomaly detector with a novel Root Cause Analysis (RCA) method based on a causal graph. This specific combination of methods allows causally justified, explainable and counterfactual RCA. The developed algorithm was applied to a simulated gripping process using robotic arms. It found the two root causes of the detected anomalies in the simulated scenarios. Josephine Rehak, Anouk Sommer, Maximilian Becker, Julius Pfrommer, Jürgen Beyerer |
INDIN | 5 |
| 2023 | Cooperative Automated Driving for Bottleneck Scenarios in Mixed TrafficabstractConnected automated vehicles (CAV), which incorporate vehicle-to-vehicle (V2V) communication into their motion planning, are expected to provide a wide range of benefits for individual and overall traffic flow. A frequent constraint or required precondition is that compatible CAVs must already be available in traffic at high penetration rates. Achieving such penetration rates incrementally before providing ample benefits for users presents a chicken-and-egg problem that is common in connected driving development. Based on the example of a cooperative driving function for bottleneck traffic flows (e.g. at a roadblock), we illustrate how such an evolutionary, incremental introduction can be achieved under transparent assumptions and objectives. To this end, we analyze the challenge from the perspectives of automation technology, traffic flow, human factors and market, and present a principle that 1) accounts for individual requirements from each domain; 2) provides benefits for any penetration rate of compatible CAVs between 0 % and 100 % as well as upward-compatibility for expected future developments in traffic; 3) can strictly limit the negative effects of cooperation for any participant and 4) can be implemented with close-to-market technology. We discuss the technical implementation as well as the effect on traffic flow over a wide parameter spectrum for human and technical aspects. Marvin V. Baumann, Jürgen Beyerer, H. Sebastian Buck, Barbara Deml, Sofie Ehrhardt, Christian Frese, D. Kleiser, Martin Lauer, Masoud Roschani, Miriam Ruf, Christoph Stiller, Peter Vortisch, Jens R. Ziehn |
IV | 2 |
| 2023 | Effects of Architectures on Continual Semantic SegmentationabstractResearch in the field of Continual Semantic Segmentation is mainly investigating novel learning algorithms to overcome catastrophic forgetting of neural networks. Most recent publications have focused on improving learning algorithms without distinguishing effects caused by the choice of neural architecture. Therefore, we study how the choice of neural network architecture affects catastrophic forgetting in class- and domain-incremental semantic segmentation. Specifically, we compare the well-researched CNNs to recently proposed Transformers and Hybrid architectures, as well as the impact of the choice of novel normalization layers and different decoder heads. We find that traditional CNNs like ResNet have high plasticity but low stability, while transformer architectures are much more stable. When the inductive biases of CNN architectures are combined with transformers in hybrid architectures, it leads to higher plasticity and stability of the model. The stability of these models can be explained by their ability to learn general features which are robust against distribution shifts. Experiments with different normalization layers show that Continual Normalization achieves the best trade-off in terms of adaptability and stability of the model, especially in domain-incremental learning. Our experiments suggest that the right choice of architecture can significantly reduce forgetting even with naive fine-tuning and confirm that for real-world applications, the architecture is an important factor in designing a continual learning model. Tobias Kalb, Niket Ahuja, Jingxing Zhou, Jürgen Beyerer |
IV | 4 |
| 2023 | Literature Review on Maneuver-Based Scenario Description for Automated Driving SimulationsabstractThe increasing complexity of automated driving functions and their growing operational design domains imply more demanding requirements on their validation. Classical methods such as field tests or formal analyses are not sufficient anymore and need to be complemented by simulations. For simulations, the standard approach is scenario-based testing, as opposed to distance-based testing primarily performed in field tests. Currently, the time evolution of specific scenarios is mainly described using trajectories, which limit or at least hamper generalizations towards variations. As an alternative, maneuver-based approaches have been proposed. We shed light on the state of the art and available foundations for this new method through a literature review of early and recent works related to maneuver-based scenario description. It includes related modeling approaches originally developed for other applications. Current limitations and research gaps are identified. Nicole Neis, Jürgen Beyerer |
IV | 2 |
| 2023 | Corner Cases in Data-Driven Automated Driving: Definitions, Properties and SolutionsabstractThe field of validation and artificial intelligence (AI) for automated driving has been a rapidly emerging field of research and development in the last few years. Despite the enormous success of machine learning (ML) in perception and robotics, the capability of ML-supported automated driving functions remains to be proven in complex real-world scenarios. Due to stringent regulations and safety concerns, it is crucial to not only be able to identify critical driving events, the corner cases, but also to eliminate them in advance by systematic and provable processes. In contrast to previous work, we analyze and systematize the causes of corner cases from the perspective of neural network interpretation, and consider the network’s performance and robustness in relation to the availability of data points used during development and validation. Moreover, we demonstrate the proposed taxonomy of corner cases on real data from multiple sensor input sources, including images and LiDAR point clouds, showing relevant properties of various corner cases. Furthermore, we discuss the possible solutions dealing with previously unknown classes and driving environments as required in future automated driving use cases. Jingxing Zhou, Jürgen Beyerer |
IV | 2 |
| 2023 | Towards Discriminative and Transferable One-Stage Few-Shot Object DetectorsabstractRecent object detection models require large amounts of annotated data for training a new classes of objects. Few-shot object detection (FSOD) aims to address this problem by learning novel classes given only a few samples. While competitive results have been achieved using two-stage FSOD detectors, typically one-stage FSODs under-perform compared to them. We make the observation that the large gap in performance between two-stage and one-stage FSODs are mainly due to their weak discriminability, which is explained by a small post-fusion receptive field and a small number of foreground samples in the loss function. To address these limitations, we propose the Few-shot RetinaNet (FSRN) that consists of: a multi-way support training strategy to augment the number of foreground samples for dense meta-detectors, an early multi-level feature fusion providing a wide receptive field that covers the whole anchor area and two augmentation techniques on query and source images to enhance transferability. Extensive experiments show that the proposed approach addresses the limitations and boosts both discriminability and transferability. FSRN is almost two times faster than two-stage FSODs while remaining competitive in accuracy, and it outperforms the state-of-the-art of one-stage meta-detectors and also some two-stage FSODs on the MS-COCO and PASCAL VOC benchmarks. Karim Guirguis, Mohamed Abdelsamad, George Eskandar, Ahmed Hendawy, Matthias Kayser, Bin Yang 0009, Jürgen Beyerer |
WACV | 7 |
| 2023 | UPAR: Unified Pedestrian Attribute Recognition and Person RetrievalabstractRecognizing soft-biometric pedestrian attributes is essential in video surveillance and fashion retrieval. Recent works show promising results on single datasets. Nevertheless, the generalization ability of these methods under different attribute distributions, viewpoints, varying illumination, and low resolutions remains rarely understood due to strong biases and varying attributes in current datasets. To close this gap and support a systematic investigation, we present UPAR, the Unified Person Attribute Recognition Dataset. It is based on four well-known person attribute recognition datasets: PA100K, PETA, RAPv2, and Market1501. We unify those datasets by providing 3,3M additional annotations to harmonize 40 important binary attributes over 12 attribute categories across the datasets. We thus enable research on generalizable pedestrian attribute recognition as well as attribute-based person retrieval for the first time. Due to the vast variance of the image distribution, pedestrian pose, scale, and occlusion, existing approaches are greatly challenged both in terms of accuracy and efficiency. Furthermore, we develop a strong baseline for PAR and attribute-based person retrieval based on a thorough analysis of regularization methods. Our models achieve state-of-the-art performance in cross-domain and specialization settings on PA100k, PETA, RAPv2, Market1501-Attributes, and UPAR. We believe UPAR and our strong baseline will contribute to the artificial intelligence community and promote research on large-scale, generalizable attribute recognition systems. The dataset is available here: https://github.com/speckean/upar_dataset Andreas Specker, Mickael Cormier, Jürgen Beyerer |
WACV | 3 |
| 2023 | Sim2real Transfer Learning for Point Cloud Segmentation: An Industrial Application Case on Autonomous DisassemblyabstractOn robotics computer vision tasks, generating and annotating large amounts of data from real-world for the use of deep learning-based approaches is often difficult or even impossible. A common strategy for solving this problem is to apply simulation-to-reality (sim2real) approaches with the help of simulated scenes. While the majority of current robotics vision sim2real work focuses on image data, we present an industrial application case that uses sim2real transfer learning for point cloud data. We provide insights on how to generate and process synthetic point cloud data in order to achieve better performance when the learned model is transferred to real-world data. The issue of imbalanced learning is investigated using multiple strategies. A novel patch-based attention network is proposed additionally to tackle this problem. Chengzhi Wu, Xuelei Bi, Julius Pfrommer, Alexander Cebulla, Simon Mangold, Jürgen Beyerer |
WACV | 6 |
| 2023 | Anticipative Feature Fusion Transformer for Multi-Modal Action AnticipationabstractAlthough human action anticipation is a task which is inherently multi-modal, state-of-the-art methods on well known action anticipation datasets leverage this data by applying ensemble methods and averaging scores of uni-modal anticipation networks. In this work we introduce transformer based modality fusion techniques, which unify multi-modal data at an early stage. Our Anticipative Feature Fusion Transformer (AFFT) proves to be superior to popular score fusion approaches and presents state-of-the-art results outperforming previous methods on EpicKitchens-100 and EGTEA Gaze+. Our model is easily extensible and allows for adding new modalities without architectural changes. Consequently, we extracted audio features on EpicKitchens-100 which we add to the set of commonly used features in the community.1 Zeyun Zhong, David Schneider 0006, Michael Voit, Rainer Stiefelhagen, Jürgen Beyerer |
WACV | 5 |
| 2022 | Causes of Catastrophic Forgetting in Class-Incremental Semantic Segmentation
Tobias Kalb, Jürgen Beyerer |
ACCV (7) | 2 |
| 2022 | For the Sake of Privacy: Skeleton-Based Salient Behavior RecognitionabstractAuthorities as well as emergency and rescue services have an increasing interest in smart support systems to ensure public safety which includes in particular behavioral analysis of pedestrians by using video surveillance systems. In order to accommodate concerns of citizens regarding their personal rights, the demand for data privacy friendly approaches, using as few information as possible, arises. In this paper, we examine existing approaches tackling the recognition of anomalous or salient behavior based solely on person pose information within the context of real-world surveillance applications. Particularly, we chose two existing state-of-the-art approaches and evaluate them on two public and an internal dataset in order to examine the overall performance of these methods for the desired task. Furthermore, we present our own approach achieving comparable results to these methods. Finally, we extend the aforementioned methods with a memory extension for modeling normal behavior, which yields on average a 4.3% higher recognition performance. Thomas Golda, Johanna Thiemich, Mickael Cormier, Jürgen Beyerer |
ICIP | 4 |
| 2022 | Towards a Better Understanding of Machine Learning based Network Intrusion Detection Systems in Industrial NetworksabstractIt is crucial in an industrial network to understand how and why a intrusion detection system detects, classifies, and reports intrusions. With the ongoing introduction of machine learning into the research area of intrusion detection, this understanding gets even more important since the used systems often appear as a black-box for the user and are no longer understandable in an intuitive and comprehensible way. We propose a novel approach to understand the internal characteristics of a machine learning based network intrusion detection system. This approach includes methods to understand which data sources the system uses, to evaluate whether the system uses linear or non-linear classification approaches, and to find out which underlying machine learning model is implemented in the system. Our evaluation on two publicly available industrial datasets shows that the detection of the data source and the differentiation between linear and non-linear models is possible with our approach. In addition, the identification of the underlying machine learning model can be accomplished with statistical significance for non-linear models. The information made accessible by our approach helps to develop a deeper understanding of the functioning of a network intrusion detection system, and contributes towards developing transparent machine learning based intrusion detection approaches. Anne Borcherding, Lukas Feldmann, Markus Karch, Ankush Meshram, Jürgen Beyerer |
ICISSP | 5 |
| 2022 | Cluster Crash: Learning from Recent Vulnerabilities in Communication StacksabstractTo ensure functionality and security of network stacks in industrial device, thorough testing is necessary. This includes blackbox network fuzzing, where fields in network packets are filled with unexpected values to test the device’s behavior in edge cases. Due to resource constraints, the tests need to be efficient and such the input values need to be chosen intelligently. Previous solutions use heuristics based on vague knowledge from previous projects to make these decisions. We aim to structure existing knowledge by defining Vulnerabil- ity Anti-Patterns for network communication stacks based on an analysis of the recent vulnerability groups Ripple20, Amnesia:33, and Urgent/11. For our evaluation, we implement fuzzing test scripts based on the Vulnerability Anti-Patterns and run them against 8 industrial device from 5 different device classes. We show (I) that similar vulnerabilities occur in implementations of the same protocol as well as in different protocols, (II) that similar vulnerabilities also spread over different device classes, and (III) that test scripts based on the Vulnerability Anti-Patterns help to identify these vulnerabilities. Anne Borcherding, Philipp Takacs, Jürgen Beyerer |
ICISSP | 3 |
| 2022 | RangeBird: Multi View Panoptic Segmentation of 3D Point Clouds with Neighborhood AttentionabstractPanoptic segmentation of point clouds is one of the key challenges of 3D scene understanding, requiring the simultaneous prediction of semantics and object instances. Tasks like autonomous driving strongly depend on these information to get a holistic understanding of their 3D environment. This work presents a novel proposal free framework for lidar-based panoptic segmentation, which exploits three different point cloud representations, leveraging their strengths and compensating their weaknesses. The efficient projection-based range view and bird's eye view are combined and further extended by a point-based network with a novel attention-based neighborhood aggregation for improved semantic features. Cluster-based object recognition in bird's eye view enables an efficient and high-quality instance segmentation. Semantic and instance segmentation are fused and further refined by a novel instance classification for the final panoptic segmentation. The results on two challenging large-scale datasets, nuScenes and SemanticKITTI, show the success of the proposed framework, which outperforms all existing approaches on nuScenes and achieves state-of-the-art results on SemanticKITTI. Fabian Duerr, Hendrik Weigel, Jürgen Beyerer |
ICRA | 3 |
| 2022 | Deep Sensor Fusion with Pyramid Fusion Networks for 3D Semantic SegmentationabstractRobust environment perception for autonomous vehicles is a tremendous challenge, which makes a diverse sensor set with e.g. camera, lidar and radar crucial. In the process of understanding the recorded sensor data, 3D semantic segmentation plays an important role. Therefore, this work presents a pyramid-based deep fusion architecture for lidar and camera to improve 3D semantic segmentation of traffic scenes. Individual sensor backbones extract feature maps of camera images and lidar point clouds. A novel Pyramid Fusion Backbone fuses these feature maps at different scales and combines the multimodal features in a feature pyramid to compute valuable multimodal, multi-scale features. The Pyramid Fusion Head aggregates these pyramid features and further refines them in a late fusion step, incorporating the final features of the sensor backbones. The approach is evaluated on two challenging outdoor datasets and different fusion strategies and setups are investigated. It outperforms recent range view based lidar approaches as well as all so far proposed fusion strategies and architectures. Hannah Schieber, Fabian Duerr, Torsten Schoen, Jürgen Beyerer |
IV | 4 |
| 2022 | Impacts of Data Anonymization on Semantic SegmentationabstractFor the development of machine learning-based driver assistance systems and highly automated driving functions, training data play a significant role in ensuring machine learning algorithms generalize well on real driving scenarios. However, data protection regulations in Europe require that individuals’ data should be processed in such a way that the individual cannot be identified from the collected data. Therefore, before camera images taken from test vehicles save on a server, license plates and faces of individuals should be anonymized first. Nevertheless, the impact of using anonymized data on the performance of machine learning algorithms remains unclear. Our work aims to evaluate the impact of anonymization on the task of semantic segmentation using diverse neural network architectures, a range of input image resolutions, and different anonymization patterns. We observe statistically significant effects of anonymizing image data on model performance and investigate methods for mitigating segmentation precision loss. Jingxing Zhou, Jürgen Beyerer |
IV | 2 |
| 2022 | Towards Heterogeneous Remote Attestation Protocolsabstract586 Paul Georg Wagner, Jürgen Beyerer |
SECRYPT | 2 |
| 2021 | Multi-Pedestrian Tracking with ClustersabstractOne of the biggest challenges in multi-pedestrian tracking arises in crowds, where missing detections can lead to wrong track-detection assignments, especially under heavy occlusion. In order to identify such situations, we cluster tracks and detections based on their overlaps and introduce different cluster states depending on the number of detections and tracks in a cluster. On the basis of this strategy, we make the following contributions. First, we propose a cluster-aware non-maximum suppression (CA-NMS) that leverages temporal information from tracks applying an increased IoU threshold in clusters with severe occlusion to reduce the number of missed detections, while at the same time limiting the number of duplicate detections. Second, for clusters with very high overlaps where detections are missing even with the CA-NMS, we utilize past track information to correct wrong assignments when missed targets are re-detected after occlusion. Furthermore, we propose a new tracking pipeline that combines the paradigms of tracking-by-detection and regression-based tracking to improve the association performance in crowded scenes. Putting all together, our tracker achieves competitive results w.r.t. the state-of-the-art on three multi-pedestrian tracking benchmarks. Our framework is analyzed with extensive ablative experiments and the impact of the proposed tracking components on the performance is evaluated. Daniel Stadler, Jürgen Beyerer |
AVSS | 2 |
| 2021 | On the Performance of Crowd-Specific Detectors in Multi-Pedestrian TrackingabstractIn recent years, several methods and datasets have been proposed to push the performance of pedestrian detection in crowded scenarios. In this study, three crowd-specific detectors are combined with a general tracking-by-detection approach to evaluate their applicability in multi-pedestrian tracking. Investigating the relation between detection and tracking accuracy, we make the interesting observation that in spite of a high detection capability, the performance in tracking can be poor and analyze the reasons behind that. However, one of the examined approaches can significantly boost the tracking performance on two benchmarks under different training configurations. It is shown that combining crowd-specific detectors with a simple tracking pipeline can achieve promising results, especially in challenging scenes with heavy occlusion. Although our tracker only relies on motion cues and no visual information is considered, applying the strong detections from the crowd-specific model, state-of-the-art results on the challenging MOT17 and MOT20 benchmarks are obtained. Daniel Stadler, Jürgen Beyerer |
AVSS | 2 |
| 2021 | Improving Multiple Pedestrian Tracking by Track Management and Occlusion HandlingabstractMulti-pedestrian trackers perform well when targets are clearly visible making the association task quite easy. However, when heavy occlusions are present, a mechanism to re-identify persons is needed. The common approach is to extract visual features from new detections and compare them with the features of previously found tracks. Since those detections can have substantial overlaps with nearby targets – especially in crowded scenarios – the extracted features are insufficient for a reliable re-identification. In contrast, we propose a novel occlusion handling strategy that explicitly models the relation between occluding and occluded tracks outperforming the feature-based approach, while not depending on a separate re-identification network. Furthermore, we improve the track management of a regression-based method in order to bypass missing detections and to deal with tracks leaving the scene at the border of the image. Finally, we apply our tracker in both temporal directions and merge tracklets belonging to the same target, which further enhances the performance. We demonstrate the effectiveness of our tracking components with ablative experiments and surpass the state-of-the-art methods on the three popular pedestrian tracking benchmarks MOT16, MOT17, and MOT20. Daniel Stadler, Jürgen Beyerer |
CVPR | 2 |
| 2021 | Improving Attribute-Based Person Retrieval By Using A Calibrated, Weighted, And Distribution-Based Distance MetricabstractTypically, person re-identification systems use so-called query images of a person to find occurrences of a person in surveillance footage. In real-world scenarios, however, often only witness descriptions and no images of a person-of-interest are available. In such cases, attribute-based retrieval can be performed based on pedestrian attribute recognition approaches. Current methods rely on calculating the Euclidean distance between query attributes and attribute predictions of gallery samples to compute ranking result lists. However, these methods do not consider the output distributions of the attribute classifier. We propose to do so by introducing a distance computation that remedies the effects of unbalanced distributions of attribute predictions. Moreover, we adapt a calibration technique to reduce the negative influences further. We also propose to weight the attributes during retrieval based on their prediction errors. In total, our approach surpasses the retrieval performance achieved by the Euclidean distance by a large margin on several datasets. In figures, we were able to increase the mAP on the RAP-2.0 and Market-1501 datasets by 4.2% and 2.4% points, respectively. Andreas Specker, Jürgen Beyerer |
ICIP | 2 |
| 2021 | Efficient Semantic Representation of Network Access Control Configuration for Ontology-based Security AnalysisabstractAssessing countermeasures and the sufficiency of security-relevant configurations within networked system architectures is a very complex task. Even the configuration of single network access control (NAC) instances can be too complex to analyse manually. Therefore, model-based approaches have manifested themselves as a solution for computer-aided configuration analysis. Unfortunately, current approaches suffer from various issues like coping with configuration-language heterogeneity or the analysis of multiple NAC instances as one overall system configuration, which is the case for the maturity of analysis goals. In this paper, we show how deriving and modelling NAC configurations’ effects solves the majority of these issues by allowing generic and simplified security analysis and model extension. The paper further presents the underlying modelling strategy to create such configuration effect representations (hereafter referred to as effective configuration) and explains how analyses based on previous approaches can still be performed. Moreover, the linking between rule representations and effective configuration is demonstrated, which enables the tracing of issues, found in the effective configuration, back to specific rules. Copyright © 2021 by SCITEPRESS – Science and Technology Publications, Lda. All rights reserved Florian Patzer, Jürgen Beyerer |
ICISSP | 2 |
| 2021 | Continual Learning for Class- and Domain-Incremental Semantic SegmentationabstractThe field of continual deep learning is an emerging field and a lot of progress has been made. However, concurrently most of the approaches are only tested on the task of image classification, which is not relevant in the field of intelligent vehicles. Only recently approaches for class-incremental semantic segmentation were proposed. However, all of those approaches are based on some form of knowledge distillation. At the moment there are no investigations on replay-based approaches that are commonly used for object recognition in a continual setting. At the same time while unsupervised domain adaption for semantic segmentation gained a lot of traction, investigations regarding domain-incremental learning in an continual setting is not well-studied. Therefore, the goal of our work is to evaluate and adapt established solutions for continual object recognition to the task of semantic segmentation and to provide baseline methods and evaluation protocols for the task of continual semantic segmentation. We firstly introduce evaluation protocols for the class- and domain-incremental segmentation and analyze selected approaches. We show that the nature of the task of semantic segmentation changes which methods are most effective in mitigating forgetting compared to image classification. Especially, in class-incremental learning knowledge distillation proves to be a vital tool, whereas in domain-incremental learning replay methods are the most effective method. Tobias Kalb, Masoud Roschani, Miriam Ruf, Jürgen Beyerer |
IV | 4 |
| 2020 | LiDAR-based Recurrent 3D Semantic Segmentation with Temporal Memory AlignmentabstractUnderstanding and interpreting a 3d environment is a key challenge for autonomous vehicles. Semantic segmentation of 3d point clouds combines 3d information with semantics and thereby provides a valuable contribution to this task. In many real-world applications, point clouds are generated by lidar sensors in a consecutive fashion. Working with a time series instead of single and independent frames enables the exploitation of temporal information. We therefore propose a recurrent segmentation architecture (RNN), which takes a single range image frame as input and exploits recursively aggregated temporal information. An alignment strategy, which we call Temporal Memory Alignment, uses ego motion to temporally align the memory between consecutive frames in feature space. A Residual Network and ConvGRU are investigated for the memory update. We demonstrate the benefits of the presented approach on two large-scale datasets and compare it to several state of-the-art methods. Our approach ranks first on the SemanticKITTI [4] multiple scan benchmark and achieves state-of-the-art performance on the single scan benchmark. In addition, the evaluation shows that the exploitation of temporal information significantly improves segmentation results compared to a single frame approach. Fabian Duerr, Mario Pfaller, Hendrik Weigel, Jürgen Beyerer |
3DV | 4 |
| 2020 | Portable Trust Anchor for OPC UA Using Auto-ConfigurationabstractWith increasing connectivity between industrial devices their attack surface grows. Consequently, secure setups have to allow these devices to distinguish trustworthy and untrustworthy communication partners. One of the most significant and wide-spread protocols for Ethernet-based data exchange between industrial devices and controls is the Open Platform Communications Unified Architecture (OPC UA). Although the OPC UA standard includes certificate-based security measures, it lacks of applicable solutions for bootstrapping trust. For secure communication, each OPC UA application is supposed to hold an application certificate which can be managed by a Global Discovery Server (GDS). Simply put, applications request their certificate from the GDS as well as information about trustworthy and revoked certificates of third parties. However, the OPC UA specifications do not suggest a secure method to establish the initial trust for the communication between an application and the GDS. Moreover, in current implementations, the administrator has to manually interchange the certificates between the peers to build sufficient trust relationships. This paper proposes an evaluated portable trust-anchor-based concept to establish this initial trust and demonstrates it solely based on standardized OPC UA communication. David Meier, Florian Patzer, Matthias Drexler, Jürgen Beyerer |
ETFA | 4 |
| 2020 | An Evaluation Of Design Choices For Pedestrian Attribute Recognition In VideoabstractPerson attribute recognition in surveillance data is a challenging task. Attributes are often visible in very localized regions and recognition thus suffers from poor image quality, changing lighting conditions, viewing angles, and occlusions. Previous research has focused predominantly on recognition in single images. In this work, we investigate the applicability of several recent strategies to include temporal information into the recognition process. We identify the most promising building blocks and create a strong baseline model, which achieves state-of-the-art attribute recognition accuracy in videos and provides a good basis for future research. Finally, we show that the resulting attributes can serve as a basis for description-based person retrieval. Andreas Specker, Arne Schumann, Jürgen Beyerer |
ICIP | 3 |
| 2020 | A Meta Model for a Comprehensive Description of Network Protocols Improving Security Testsabstract671 Steffen Pfrang, David Meier, Andreas Fleig, Jürgen Beyerer |
ICISSP | 4 |
| 2019 | Human Pose Estimation for Real-World Crowded ScenariosabstractHuman pose estimation has recently made significant progress with the adoption of deep convolutional neural networks and many applications have attracted tremendous interest in recent years. However, many of these applications require pose estimation for human crowds, which still is a rarely addressed problem. For this purpose this work explores methods to optimize pose estimation for human crowds, focusing on challenges introduced with larger scale crowds like people in close proximity to each other, mutual occlusions, and partial visibility of people due to the environment. In order to address these challenges, multiple approaches are evaluated including: the explicit detection of occluded body parts, a data augmentation method to generate occlusions and the use of the synthetic generated dataset JTA [3]. In order to overcome the transfer gap of JTA originating from a low pose variety and less dense crowds, an extension dataset is created to ease the use for real-world applications. Thomas Golda, Tobias Kalb, Arne Schumann, Jürgen Beyerer |
AVSS | 4 |
| 2019 | Suggesting Gaze-based Selection for Surveillance ApplicationsabstractThe selection operation is a basic input operation when interacting with a computer. Traditional manual selection methods like mouse input are challenging when interacting with a dynamic scene containing moving objects as it occurs in surveillance applications. In this contribution, we give an overview on gaze-based selection as a fast and intuitive alternative method considering typical selection tasks in surveillance applications. In this context, we report the results of an evaluation on initialization of an object-tracking algorithm performed by eighteen expert video analysts. Besides its benefits, gaze-based selection is difficult if selection objects are small and close together. We provide first results of a pilot study evaluating a distance measure, which might have the potential of making gaze-based selection more robust. Finally, to leverage gaze-based selection becoming a common technique, low-cost eye-tracking devices have to be available which achieve the same accuracy as the high-end devices. As such cheap eye-trackers only recently became available, it is so far not evident, which accuracy we can expect. Hence, we report the results of an accuracy evaluation of the Tobii 4C conducted with twelve students performing a calibration-style task. Jutta Hild, Elisabeth Peinsipp-Byma, Michael Voit, Jürgen Beyerer |
AVSS | 4 |
| 2019 | An Interactive Framework for Cross-modal Attribute-based Person RetrievalabstractPerson re-identification systems generally rely on a query person image to find additional occurrences of this person across a camera network. In many real-world situations, however, no such query image is available and witness testimony is the only clue upon which to base a search. Cross-modal re-identification based on attribute queries can help in such cases but currently yields a low matching accuracy which is often not sufficient for practical applications. In this work we propose an interactive feedback-driven framework, which successfully bridges the modality gap and achieves a significant increase in accuracy by 47% in mean average precision (mAP) compared to the fully automatic cross-modal state-of-the-art. We further propose a cluster-based feedback method as part of the framework, which outperforms naïve user feedback by more than 9% mAP. Our results set a new state-of-the-art for fully automatic and feedback-driven cross-modal attribute-based re-identification on two public datasets. Andreas Specker, Arne Schumann, Jürgen Beyerer |
AVSS | 3 |
| 2019 | The Industrie 4.0 Asset Administration Shell as Information Source for Security AnalysisabstractOne of the essential concepts of the Reference Architecture Model Industrie 4.0 (RAMI4.0) is the uniform modelling of assets by means of a common meta-data model called the Asset Administration Shell (AAS). However, important practical experience with this concept is still missing, as not many use cases for the AAS have yet been implemented. Thus, practical issues within the AAS concept and respective solutions are hard to identify. In this paper, presents our experience with the implementation of an AAS use case. The AAS is used as information source to create an ontology, which is then used for security analysis. The paper discusses the use-case-specific modelling language selection and provides a practical examination of several of our implementations that use OWL and OPC UA together. Furthermore, it provides recommendations for the implementation of Asset Administration Shells for this and similar use cases. Florian Patzer, Friedrich Volz, Thomas Usländer, Immanuel Blöcher, Jürgen Beyerer |
ETFA | 5 |
| 2019 | Design of an Example Network Protocol for Security Tests Targeting Industrial Automation SystemsabstractEmerging concepts like Industrial Internet of Things (IIOT) and Industrie 4.0 require Industrial Automation and Control Systems (IACS) to be connected via networks and even to the Internet. These connections raise the importance of security for those devices enormously. Security testing for IACS aims at searching for vulnerabilities which can be utilized by attackers from the network. Once discovered, those gaps should be closed with patches before they can get exploited. Different tools utilized for this kind of security testing are dealing with network protocols. In practice, they suffer from peculiarities being present in common industrial automation protocols like OPC UA and Profinet IO. This paper tries to improve the situation by providing an extensive overview of network packet structures and network protocol behavior. Based on this analysis, an example protocol has been developed. The idea behind this artificial network protocol is that tools which are able to handle all the specialties of this protocol, are able to handle every imaginable protocol. Finally, those tools can be used to conduct exhaustive security tests for IACS. Steffen Pfrang, Mark Giraud, Anne Borcherding, David Meier, Jürgen Beyerer |
ICISSP | 5 |
| 2019 | Towards Real-Time Detection and Mitigation of Driver Frustration using SVMabstractDriving in stressful and frustrating situations remains a common issue in daily traffic scenarios and has been shown to increase the risk for hazardous and aggressive driving style. Aiming to improve road safety with intelligent systems, frustration has to be detected continuously as well as robust and mitigation strategies must be applied effectively. Since both, the modeling of frustration over time as well as the design and timing of applications for frustration mitigation, are complex tasks, we divided this work in two parts: (1) A driving simulator experiment was conducted to collect a dataset and to validate a driving context related frustration induction method. With this dataset we developed a bimodal frustration detection for the driving context using a temporal support vector machine. The detection combines drivers' visual facial features with heart rate measurements and yields an accuracy of 88.7% (AUC of ROC). (2) We applied the real-time frustration detection in a second simulator study to evaluate two application scenarios for frustration mitigation, which include a frustration sensitive ambient light and an autonomous driving assistant. Both applications were examined for their frustration mitigating effect as well as in terms of user experience. The results provide a helpful basis to develop future intelligent frustration mitigation systems. Sebastian Zepf, Tobias Stracke, Alexander Schmitt, Florian van de Camp, Jürgen Beyerer |
ICMLA | 5 |
| 2019 | Patch-based facial texture super-resolution by fitting 3D face models
Chengchao Qu, Eduardo Monari, Tobias Schuchert, Jürgen Beyerer |
Mach. Vis. Appl. | 4 |
| 2019 | Comprehensive Analysis of Deep Learning-Based Vehicle Detection in Aerial ImagesabstractVehicle detection in aerial images is a crucial image processing step for many applications such as screening of large areas as used for surveillance, reconnaissance, or rescue tasks. In recent years, several deep learning-based frameworks have been proposed for object detection. However, these detectors were developed for data sets that considerably differ from aerial images. In this paper, we systematically investigate the potential of fast R-CNN and faster R-CNN for aerial images, which achieve top performing results on common detection benchmark data sets. Therefore, the applicability of eight state-of-the-art object proposal methods used to generate a set of candidate regions and of both detectors is examined. Relevant adaptations to account for the characteristics of the aerial images are provided. To overcome the shortcomings of the original approach in the case of handling small instances, we further propose our own networks that clearly outperform state-of-the-art methods for vehicle detection in aerial images. Furthermore, we analyze the impact of the different adaptations with respect to various ground sampling distances to provide a guideline for detecting small objects in aerial images. All experiments are performed on two publicly available data sets to account for differing characteristics such as varying object sizes, number of objects per image, and varying backgrounds. Lars Wilko Sommer, Tobias Schuchert, Jürgen Beyerer |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Comprehensive Evaluation of Deep Learning based Detection Methods for Vehicle Detection in Aerial ImageryabstractAutomatic analysis of aerial imagery acquired by satellites, planes and UAVs facilitates several applications such as traffic monitoring, surveillance, search and rescue tasks. These applications have in common the need for an accurate object detection. In recent years, applying Faster R-CNN, a deep learning based detection method, outperformed conventional detection methods for the task of vehicle detection in aerial imagery. For this, adaptations to the characteristics of aerial imagery are necessary. In this paper, we adapt several state-of-the-art detectors including Faster R-CNN, SSD, and YOLOv2 and analyze the detection performance with respect to object categories, ground sampling distances and inference time. Furthermore, we examine the impact of adding more semantic information for each detector separately. For that purpose, we propose an extension of YOLOv2 by adding a deconvolutional module. We achieve state-of-the-art results on the publicly available DOTA dataset. Oliver Acatay, Lars Wilko Sommer, Arne Schumann, Jürgen Beyerer |
AVSS | 4 |
| 2018 | Attribute-based Person Retrieval and Search in Video SequencesabstractThe search for persons based on their visual appearance is an important task of modern surveillance systems which can be supported by automatic person re-identification approaches. However, such approaches are generally image-based and thus require a query image as input. In cases where only a witness description is available the task turns into a cross-modal text-to-image search problem which requires specialized approaches. In this work we describe an approach for person search in video data based purely on attribute witness descriptions. We first develop an ensemble of classifiers for robust attribute classification. We then extend the approach to full person search by combining it with a person detector. Given an initial high-confidence match in the video we use temporal information to explore and return full person tracks. We evaluate our approach on the AVSS 2018 Soft Biometric Retrieval Challenge dataset. Our approach manages to find the correct person at rank-1 in 71.79% of all cases. Arne Schumann, Andreas Specker, Jürgen Beyerer |
AVSS | 3 |
| 2018 | Onthology-based Masking Loss for Improved Generalization in Remote Sensing Semantic Image RetrievalabstractSemantic image retrieval can significantly reduce the time required to process the vast amounts of remote sensing image data and support applications such as detection of illegal fishing or logging or analysis of growth and change in residential and industrial areas. In this work we propose a novel method for remote sensing semantic retrieval based on a masking loss for convolutional neural networks which allows multi-dataset training despite different and incompatible semantic classes in different datasets. We achieve improved generalization when the trained model is applied in a realistic cross-dataset setting. In addition to this we perform a thorough evaluation of several design choices which are popular in other retrieval tasks, most notably the impact of specialized ranking losses, and formulate guidelines for future research. Our trained models are evaluated on the recent PatternNet dataset and the established WHU-RS19 and UCM dataset. We outperform the state-of-the-art on PatternNet, UCM, and WHU-RS19 by 29.3%, 17.2%, and 4.5%, respectively. Arne Schumann, Lars Wilko Sommer, Max Vogler, Jürgen Beyerer |
AVSS | 4 |
| 2018 | Ensemble of Two-Stage Regression Based Detectors for Accurate Vehicle Detection in Traffic Surveillance DataabstractThe growing amount of traffic surveillance data results in an increased need for automatic detection systems to analyze the data. For this purpose, deep learning based detection frameworks like Faster R-CNN and SSD have been employed in recent years. Though the detection accuracy is clearly improved compared to conventional detection methods, there exists large potential for further improvements especially in case of adverse weather conditions. In this paper, we employ the RefineDet detection framework as it combines advantages of several detection frameworks including Faster R-CNN and SSD. We use an ensemble of two detectors with different base networks to generate detections that are more robust. For this, SENets - the winner of the ImageNet2017 classification challenge - are used in addition to ResNet-50. To account for small vehicles in the background and strong variation in vehicle scale, we apply multi-scale testing. Our proposed detector achieves top-performing results on the UA-DETRAC dataset especially in case of rainy and nighttime scenarios. Lars Wilko Sommer, Oliver Acatay, Arne Schumann, Jürgen Beyerer |
AVSS | 4 |
| 2018 | Comparison of Angle and Size Features with Deep Learning for Emotion Recognition
Patrick Dunau, Marco F. Huber, Jürgen Beyerer |
CIARP | 3 |
| 2018 | Towards Computer-Aided Security Life Cycle Management for Critical Industrial Control Systems
Florian Patzer, Ankush Meshram, Pascal Birnstill, Christian Haas 0005, Jürgen Beyerer |
CRITIS | 5 |
| 2018 | Predicting observer's task from eye movement patterns during motion image analysisabstractPredicting an observer's tasks from eye movements during several viewing tasks has been investigated by several authors. This contribution adds task prediction from eye movements tasks occurring during motion image analysis: Explore, Observe, Search, and Track. For this purpose, gaze data was recorded from 30 human observers viewing a motion image sequence once under each task. For task decoding, the classification methods Random Forest, LDA, and QDA were used; features were fixation- or saccade-related measures. Best accuracy for prediction of the three tasks Observe, Search, Track from the 4-minute gaze data samples was 83.7% (chance level 33%) using Random Forest. Best accuracy for prediction of all four tasks from the gaze data samples containing the first 30 seconds of viewing was 59.3% (chance level 25%) using LDA. Accuracy decreased significantly for task prediction on small gaze data chunks of 5 and 3 seconds, being 45.3% and 38.0% (chance 25%) for the four tasks, and 52.3% and 47.7% (chance 33%) for the three tasks. Jutta Hild, Michael Voit, Christian Kühnle, Jürgen Beyerer |
ETRA | 4 |
| 2018 | Search Area Reduction Fast-RCNN for Fast Vehicle Detection in Large Aerial ImageryabstractAccurate detection of objects in aerial imagery is a crucial image processing step for many applications, such as traffic monitoring, surveillance, reconnaissance and rescue tasks. Recently, conventional methods for vehicle detection in aerial imagery are outperformed by deep learning based detection frameworks like Faster R-CNN. To allow the detection of small vehicles in the range of 10x20 pixels, only shallow layers of standard models like VGG-16 provide a sufficiently high spatial resolution. However, this adaptation to the characteristics of aerial imagery results in poor inference time as proposals or detections are predicted at each feature map location. In this paper, we propose an adaptive model which reduces the input region for the detection module and consequently reduces inference time. For this, we extend Faster R-CNN by an additional Search Area Reduction module which divides the input image into regions and predicts a confidence score of how likely a region contains at least one object. Many image regions, particularly in rural areas, do not contain any vehicles and are filtered out by our approach. This significantly reduces inference time of later stages. Our proposed framework achieves state-of-the-art detection results on a publicly available dataset while the inference time of the object detection stage is reduced by more than 75%.11Our custom layers and network architecture have been made available at: https://github.com/vehicledetect/sar-frcnn. Lars Wilko Sommer, Nicole Schmidt, Arne Schumann, Jürgen Beyerer |
ICIP | 4 |
| 2018 | Advancing Protocol Fuzzing for Industrial Automation and Control SystemsabstractS.570-580 Steffen Pfrang, David Meier, Michael Friedrich 0003, Jürgen Beyerer |
ICISSP | 4 |
| 2018 | Distributed Usage Control Enforcement through Trusted Platform Modules and SGX EnclavesabstractIn the light of mobile and ubiquitous computing, sharing sensitive information across different computer systems has become an increasingly prominent practice. This development entails a demand of access control measures that can protect data even after it has been transferred to a remote computer system. In order to address this problem, sophisticated usage control models have been developed. These models include a client side reference monitor (CRM) that continuously enforces protection policies on foreign data. However, it is still unclear how such a CRM can be properly protected in a hostile environment. The user of the data on the client system can influence the client's state and has physical access to the system. Hence technical measures are required to protect the CRM on a system, which is legitimately used by potential attackers. Existing solutions utilize Trusted Platform Modules (TPMs) to solve this problem by establishing an attestable trust anchor on the client. However, the resulting protocols have several drawbacks that make them infeasible for practical use. This work proposes a reference monitor implementation that establishes trust by using TPMs along with Intel SGX enclaves. First we show how SGX enclaves can realize a subset of the existing usage control requirements. Then we add a TPM to establish and protect a powerful enforcement component on the client. Ultimately this allows us to technically enforce usage control policies on an untrusted remote system. Paul Georg Wagner, Pascal Birnstill, Jürgen Beyerer |
SACMAT | 3 |
| 2018 | Semantic Labeling Based Vehicle Detection in Aerial ImageryabstractThe increasing abundance of available aerial image and video data facilitates many applications, such as disaster relief, analysis of traffic flows, city planning, search tasks, and situation recognition. Automated detection systems that provide accurate detections of all relevant objects, e.g. vehicles, are essential for such applications. However, current detection frameworks, such as Faster R-CNN, are prone to cause false positive detections due to objects with shapes similar to vehicles, such as windows and solar panels on buildings. To address this issue, we propose two multi-task models that combine the detection task and a semantic labeling task, which induces more scene knowledge into the model. Through a shared global feature map we can improve detection results significantly. Additionally, by explicitly merging features of the semantic labeling branch into the region pooling step of the detection framework we can further reduce detection errors. We evaluate both models on the popular Potsdam dataset and outperform recent related work. Kun Nie, Lars Wilko Sommer, Arne Schumann, Jürgen Beyerer |
WACV | 4 |
| 2018 | Multi Feature Deconvolutional Faster R-CNN for Precise Vehicle Detection in Aerial ImageryabstractAccurate detection of objects in aerial images is an important task for many applications such as traffic monitoring, surveillance, reconnaissance and rescue tasks. Recently, deep learning based detection frameworks clearly improved the detection performance on aerial images compared to conventional methods comprised of hand-crafted features and a classifier within a sliding window approach. These deep learning based detection frameworks use the output of the last convolutional layer as feature map for localization and classification. Due to the small size of objects in aerial images, only shallow layers of standard models like VGG-16 or small networks are applicable in order to provide a sufficiently high feature map resolution. However, high-resolution feature maps offer less semantic and contextual information, which results in approaches being more prone to false alarms due to objects with similar shapes especially in case of tiny objects. In this paper, we extend the Faster R-CNN detection framework to cope this issue. Therefore, we apply a deconvolutional module that up-samples low-dimensional feature maps of deep layers and combines the up-sampled features with the features of shallow layers while the feature map resolution is kept sufficiently high to localize tiny objects. Our proposed deconvolutional framework clearly outperforms state-of-the-art methods on two publicly available datasets. Lars Wilko Sommer, Arne Schumann, Tobias Schuchert, Jürgen Beyerer |
WACV | 4 |
| 2017 | Drone-vs-Bird detection challenge at IEEE AVSS2017abstractSmall drones are a rising threat due to their possible misuse for illegal activities, in particular smuggling and terrorism. The project SafeShore, funded by the European Commission under the Horizon 2020 program, has launched the “drone-vs-bird detection challenge” to address one of the many technical issues arising in this context. The goal is to detect a drone appearing at some point in a video where birds may be also present: the algorithm should raise an alarm and provide a position estimate only when a drone is present, while not issuing alarms on birds. This paper reports on the challenge proposal, evaluation, and results. Angelo Coluccia, Marian Ghenescu, Tomas Piatrik, Geert De Cubber, Arne Schumann, Lars Wilko Sommer, Johannes Klatte, Tobias Schuchert, Jürgen Beyerer, Mohammad Farhadi, Ruhallah Amandi, Cemal Aker, Sinan Kalkan, Nabin Sharma, Sultan Daud Khan, Khan Makkah, Michael Blumenstein |
AVSS | 9 |
| 2017 | Deep cross-domain flying object classification for robust UAV detectionabstractRecent progress in the development of unmanned aerial vehicles (UAVs) causes serious safety issues for mass events and safety-sensitive locations like prisons or airports. To address these concerns, robust UAV detection systems are required. In this work, we propose an UAV detection framework based on video images. Depending on whether the video images are recorded by static cameras or moving cameras, we initially detect regions that are likely to contain an object by median background subtraction or a deep learning based object proposal method, respectively. Then, the detected regions are classified into UAV or distractors, such as birds, by applying a convolutional neural network (CNN) classifier. To train this classifier, we use our own dataset comprised of crawled and self-acquired drone images, as well as bird images from a publicly available dataset. We show that, even across a significant domain gap, the resulting classifier can successfully identify UAVs in our target dataset. We evaluate our UAV detection framework on six challenging video sequences that contain UAVs at different distances as well as birds and background motion. Arne Schumann, Lars Wilko Sommer, Johannes Klatte, Tobias Schuchert, Jürgen Beyerer |
AVSS | 5 |
| 2017 | Semantic labeling for improved vehicle detection in aerial imageryabstractGrowing cities and increasing traffic densities result in an increased demand for applications such as traffic monitoring, traffic analysis, and support of rescue work. These applications share the need for accurate detection of relevant vehicles, e.g. in aerial imagery. Recently, the application of deep learning based detection frameworks like Faster R-CNN clearly outperformed conventional detection methods for vehicle detection in aerial images. In this paper, we propose a detection framework that fuses Faster R-CNN and semantic labeling to integrate contextual information. We achieve an improved detection performance by decreasing the number of false positive detections while the number of candidate regions to classify is reduced. To demonstrate the generalization of our approach, we evaluate our detection framework for various ground sampling distances on a publicly available dataset. Lars Wilko Sommer, Kun Nie, Arne Schumann, Tobias Schuchert, Jürgen Beyerer |
AVSS | 5 |
| 2017 | Flying object detection for automatic UAV recognitionabstractWith the increasing use of unmanned aerial vehicles (UAVs) by consumers, automatic UAV detection systems have become increasingly important for security services. In such a system, video imagery is a core modality for the detection task, because it can cover large areas and is very cost-effective to acquire. Many detection systems consist of two parts: flying object detection and subsequent object classification. In this work, we investigate the suitability of a number of flying object detection approaches for the task of UAV detection based on video data from static and moving cameras. We compare approaches based on image differencing with object proposal detectors which are learned from data. Finally, we classify each detection by a convolutional neural network (CNN) into the classes UAV or clutter. Our approach is evaluated on six sequences of challenging real world data which contain multiple UAVs, birds, and background motion. Lars Wilko Sommer, Arne Schumann, Thomas Muller, Tobias Schuchert, Jürgen Beyerer |
AVSS | 5 |
| 2017 | Introducing remote attestation and hardware-based cryptography to OPC UAabstractIn this paper we investigate whether and how hardware-based roots of trust, namely Trusted Platform Modules (TPMs) can improve the security of the communication protocol OPC UA (Open Platform Communications Unified Architecture) under reasonable assumptions, i.e. the Dolev-Yao attacker model. Our analysis shows that TPMs may serve for generating (RNG) and securely storing cryptographic keys, as cryptocoprocessors for weak systems, as well as for remote attestation. We propose to include these TPM functions into OPC UA via so-called ConformanceUnits, which can serve as building blocks of profiles that are used by clients and servers for negotiating the parameters of a session. Eventually, we present first results regarding the performance of a client-server communication including an additional OPC UA server providing remote attestation of other OPC UA servers. Pascal Birnstill, Christian Haas 0003, Daniel Hassler, Jürgen Beyerer |
ETFA | 4 |
| 2017 | Towards the modelling of complex communication networks in AutomationMLabstractFor several decades production systems were considered as closed and decoupled units, where information and network security has not been an issue. This is changing rapidly, since in the age of smart factories, production systems and office IT are growing together so as to transform entire value chains into interconnected distributed systems. By this means, production systems inherit the security challenges of office IT networks connected over the Internet. Therefore and as they tend to be operated for a much longer period of time, a prospective design of security mechanisms is mandatory. For some time, the design process of production systems gets modernised by the Automation Markup Language (AutomationML, IEC 62714). AutomationML incorporates formats of all engineering phases of production systems, thus allowing engineers to model production systems on various levels of abstraction. The language also provides building blocks for modelling the network infrastructure, which are presented in the AutomationML Communication whitepaper. However, the level of detail that can be captured is currently not sufficient for modelling most network protocols and therefore any network security concept. Therefore, we propose an extension to the AutomationML Communication whitepaper and its best practice recommendations, which allows us to model networks according to the established ISO/OSI model. Using this extension we show that concepts like network separation can be modelled and validated. Florian Patzer, Aranya Sarkar, Pascal Birnstill, Miriam Schleipen, Jürgen Beyerer |
ETFA | 5 |
| 2017 | A Multi-agent Approach to Model and Analyze the Behavior of Vessels in the Maritime DomainabstractS.200-207 Mathias Anneken, Yvonne Fischer, Jürgen Beyerer |
ICAART (1) | 3 |
| 2017 | Unconstrained Face Detection and Open-Set Face Recognition ChallengeabstractFace detection and recognition benchmarks have shifted toward more difficult environments. The challenge presented in this paper addresses the next step in the direction of automatic detection and identification of people from outdoor surveillance cameras. While face detection has shown remarkable success in images collected from the web, surveillance cameras include more diverse occlusions, poses, weather conditions and image blur. Although face verification or closed-set face identification have surpassed human capabilities on some datasets, open-set identification is much more complex as it needs to reject both unknown identities and false accepts from the face detector. We show that unconstrained face detection can approach high detection rates albeit with moderate false accept rates. By contrast, open-set face recognition is currently weak and requires much more attention. Manuel Günther, Peiyun Hu, Christian Herrmann 0001, Chi-Ho Chan, Min Jiang 0003, Shufan Yang, Akshay Raj Dhamija, Deva Ramanan, Jürgen Beyerer, Josef Kittler, Mohamad Al Jazaery, Mohammad Iqbal Nouyed, Guodong Guo, Cezary Stankiewicz, Terrance E. Boult |
IJCB | 9 |
| 2017 | Probabilistic Surface Inference for Industrial Inspection PlanningabstractOptimized machine vision setups are a requirement for precise and efficient product quality assurance. As the design space is high-dimensional, a manual design requires a lot of engineering experience and experimental work, associated with high costs and often non-optimal results. For automatic evaluation, there are evaluation metrics such as measurement uncertainty and scan resolution to evaluate the quality of an inspection. However, it is not trivial how to combine different criteria to optimize the setup based on the inspection requirements. We propose to fuse the metrics through a probabilistic surface inference to quantify the amount of information gained by a specific setup configuration. To this end, the product surface is modeled by a random process and the problem is adapted to a Gaussian Process (GP) inference. We introduce a local inference based on the local surface orientation and propose a novel spectrum-based approach to determine the GP parameters based on the required inspection resolution. The inference results, based on simulations, are demonstrated for the inspection of a cylinder head. Mahsa Mohammadikaji, Stephan Bergmann, Stephan Irgenfried, Jürgen Beyerer, Carsten Dachsbacher, Heinz Wörn |
WACV | 4 |
| 2017 | Robust 3D Patch-Based Face HallucinationabstractIncorporating 3D information has proven to be effective in many computer vision tasks and it is no exception in the context of facial analysis. However, limited application has been witnessed in face hallucination (FH), probably due to the difficulty of fitting 3D models onto low-resolution (LR) images. This paper presents a pure 3D approach to address this problem. By extending the LR image formation process to the 3D domain, the classic Lucas-Kanade algorithm is exploited to improve the precision of the error-prone 3D model fitting on LR images. The established correspondence between the input image and 3D training textures then facilitates reconstruction of high-resolution (HR) patches directly on the mesh, which can be employed to render realistic frontal faces for recognition. Extensive evaluation on several publicly available datasets reveals superior qualitative and quantitative results over state-of-the-art methods in fitting, FH and recognition, which shows the advantage of the proposed 3D framework over its 2D rivals, especially for non-frontal head poses and low image quality. Chengchao Qu, Christian Herrmann 0001, Eduardo Monari, Tobias Schuchert, Jürgen Beyerer |
WACV | 5 |
| 2017 | Fast Deep Vehicle Detection in Aerial ImagesabstractVehicle detection in aerial images is a crucial image processing step for many applications like screening of large areas. In recent years, several deep learning based frameworks have been proposed for object detection. However, these detectors were developed for datasets that considerably differ from aerial images. In this paper, we systematically investigate the potential of Fast R-CNN and Faster R-CNN for aerial images, which achieve top performing results on common detection benchmark datasets. Therefore, the applicability of 8 state-of-the-art object proposals methods used to generate a set of candidate regions and of both detectors is examined. Relevant adaptations of the object proposals methods are provided. To overcome shortcomings of the original approach in case of handling small instances, we further propose our own network that clearly outperforms state-of-the-art methods for vehicle detection in aerial images. All experiments are performed on two publicly available datasets to account for differing characteristics such as ground sampling distance, number of objects per image and varying backgrounds. Lars Wilko Sommer, Tobias Schuchert, Jürgen Beyerer |
WACV | 3 |
| 2017 | Fast algorithm for 2D fragment assembly based on partial EMD
Shuang-Min Chen, Zhenyu Shu, Shi-Qing Xin, Jieyu Zhao 0002, Guang Jin, Rong Zhang 0007, Jürgen Beyerer |
Vis. Comput. | 8 |
| 2016 | Low-resolution Convolutional Neural Networks for video face recognitionabstractSecurity and safety applications such as surveillance or forensics demand face recognition in low-resolution video data. We propose a face recognition method based on a Convolutional Neural Network (CNN) with a manifold-based track comparison strategy for low-resolution video face recognition. The low-resolution domain is addressed by adjusting the network architecture to prevent bottlenecks or significant upscaling of face images. The CNN is trained with a combination of a large-scale self-collected video face dataset and large-scale public image face datasets resulting in about 1.4M training images. To handle large amounts of video data and for effective comparison, the CNN face descriptors are compared efficiently on track level by local patch means. Our setup achieves 80.3 percent accuracy on a 32×32 pixels low-resolution version of the YouTube Faces Database and outperforms local image descriptors as well as the state-of-the-art VGG-Face network [20] in this domain. The superior performance of the proposed method is confirmed on a self-collected in-the-wild surveillance dataset. Christian Herrmann 0001, Dieter Willersinn, Jürgen Beyerer |
AVSS | 3 |
| 2016 | Gaze-based moving target acquisition in real-time full motion videoabstractReal-time moving target acquisition in full motion video is a challenging task. Mouse input might fail if targets move fast, unpredictably, or are only visible for a short period of time. In this paper, we describe an experiment with expert video analysts (N=26) which perform moving target acquisition by selecting targets in a full motion video sequence presented on a desktop computer. The results show that using gaze input (gaze pointing + manual key press), the participants were able to perform with significantly shorter completion times than with mouse input. Error rates (represented by target misses) and acquisition precision were similar. Subjective ratings of user satisfaction resulted in similar or even better scores for the gaze interaction. Jutta Hild, Christian Kühnle, Jürgen Beyerer |
ETRA | 3 |
| 2016 | Moving target acquisition by gaze pointing and button press using hand or footabstractThis work explores gaze-based interaction for moving target acquisition. In a pilot study, three interaction techniques are compared: gaze and manual button press (gaze + hand), gaze and foot button press (gaze + foot), and traditional mouse input. In a controlled scenario using a circle acquisition paradigm, participants perform moving target acquisition for targets differing in speed, direction of motion and motion pattern. The results show similar hit rates for the three techniques. Target acquisition completion time is significantly faster for the gaze-based techniques compared to mouse input. Jutta Hild, Patrick Petersen, Jürgen Beyerer |
ETRA | 3 |
| 2016 | Iris tracking using extended object tracking
Patrick Dunau, Jürgen Beyerer |
FUSION | 2 |
| 2016 | Motion segmentation and appearance change detection based 2D hand tracking
Jan Hendrik Hammer, Michael Voit, Jürgen Beyerer |
FUSION | 3 |
| 2016 | Anomaly Detection using B-spline Control Points as Feature Space in Annotated Trajectory Data from the Maritime DomainabstractS.250-257 Mathias Anneken, Yvonne Fischer, Jürgen Beyerer |
ICAART (2) | 3 |
| 2016 | Capturing ground truth super-resolution dataabstractSuper-resolution (SR) offers an effective approach to boost quality and details of low-resolution (LR) images to obtain high-resolution (HR) images. Despite the theoretical and technical advances in the past decades, it still lacks plausible methodology to evaluate and compare different SR algorithms. The main cause to this problem lies in the missing ground truth data for SR. Unlike in many other computer vision tasks, where existing image datasets can be utilized directly, or with a little extra annotation work, evaluating SR requires that the dataset contain both LR and the corresponding HR ground truth images of the same scene captured at the same time. This work presents a novel prototype camera system to address the aforementioned difficulties of acquiring ground truth SR data. Two identical camera sensors equipped with a wide-angle lens and a telephoto lens respectively, share the same optical axis by placing a beam splitter in the optical path. The back-end program can then trigger their shutters simultaneously and precisely register the region of interests (ROIs) of the LR and HR image pairs in an automated manner free of sub-pixel interpolation. Evaluation results demonstrate the special characteristics of the captured ground truth HR-LR face images compared to the simulated ones. The dataset is made freely available for noncommercial research purposes. Chengchao Qu, Ding Luo, Eduardo Monari, Tobias Schuchert, Jürgen Beyerer |
ICIP | 5 |
| 2016 | Intervention-free selection using EEG and eye trackingabstractIn this paper, we show how recordings of gaze movements (via eye tracking) and brain activity (via electroencephalography) can be combined to provide an interface for implicit selection in a graphical user interface. This implicit selection works completely without manual intervention by the user. In our approach, we formulate implicit selection as a classification problem, describe the employed features and classification setup and introduce our experimental setup for collecting evaluation data. With a fully online-capable setup, we can achieve an F_0.2-score of up to 0.74 for temporal localization and a spatial localization accuracy of more than 0.95. Felix Putze, Johannes Popp, Jutta Hild, Jürgen Beyerer, Tanja Schultz |
ICMI | 4 |
| 2016 | Knowing when you don't: Bag of visual words with reject option for automatic visual inspection of bulk materialsabstractVisual inspection of bulk material is the thorough optical inspection of streams of granular material to assess their quality or to detect defective objects. Examples are found in mining (discovery of ores), recycling (sorting waste from reusable material) and food safety (detection of pathogens). In these applications, it is generally not feasible or even possible to provide an accurate and exhaustive training set of all the materials that can be encountered during the inspection. Instead, classification has to be performed in an open world setting, i.e., with the option to recognize and reject unknown objects. Despite the practical relevance, prior work on this topic is surprisingly sparse. Here, we present a method to augment bag of visual words object descriptors by an additional unknown word that encodes outliers. The method depends on only few parameters that have a clear interpretation and is suitable for the application in the field. We demonstrate the performance of our approach using two real-world datasets and compare it to a related method. The experiments show that our method significantly outperforms classification with a closed world assumption as well as the related method. Matthias Richter 0003, Thomas Längle, Jürgen Beyerer |
ICPR | 3 |
| 2016 | A survey on moving object detection for wide area motion imageryabstractWide Area Motion Imagery (WAMI) enables the surveillance of tens of square kilometers with one airborne sensor Each image can contain thousands of moving objects. Applications such as driver behavior analysis or traffic monitoring require precise multiple object tracking that is dependent on initial detections. However, low object resolution, dense traffic, and imprecise image alignment lead to split, merged, and missing detections. No systematic evaluation of moving object detection exists so far although many approaches have been presented in the literature. This paper provides a detailed overview of existing methods for moving object detection in WAMI data. Also we propose a novel combination of short-term background subtraction and suppression of image alignment errors by pixel neighborhood consideration. In total, eleven methods are systematically evaluated using more than 160,000 ground truth detections of the WPAFB 2009 dataset. Best performance with respect to precision and recall is achieved by the proposed one. Lars Wilko Sommer, Michael Teutsch, Tobias Schuchert, Jürgen Beyerer |
WACV | 4 |
| 2015 | Message from general chairsabstractAVSS is the premier annual international conference in the field of video and signal-based surveillance that brings together experts from academia, industry, and government to advance theories, methods, systems, and applications related to surveillance. Jürgen Beyerer, Rainer Stiefelhagen |
AVSS | 1 |
| 2015 | A user study on anonymization techniques for smart video surveillanceabstractA key mechanism of privacy-aware smart video surveillance is anonymization of video data. We conducted a user study with a response of 103 participants in order to investigate which pixel operations are suitable for protecting persons' identities while, at the same time, allowing a human operator to recognize persons' activities i.e., preserving the utility of the video data. Regarding the activities in the data set, namely stealing, fighting, and dropping a bag, our data does not approve the common hypothesis that privacy and utility of video data are necessarily trade-off. Pascal Birnstill, Daoyuan Ren, Jürgen Beyerer |
AVSS | 3 |
| 2015 | Impact of resolution and image quality on video face analysisabstractLow-resolution face analysis suffers more significantly from quality degradations than high-resolution analysis. In this work, we will investigate how several face analysis steps are influenced by low image quality and how this relates to the low resolution. In the first step, a simulation of different effects on image quality, namely low resolution, compression artifacts, motion blur and noise is performed and the impact on face detection, registration and recognition is analyzed. Depending on the situation, it becomes obvious that the low resolution is sometimes a minor degrading effect, outmatched by a single one or a combination of the further effects. When addressing real-world face recognition from surveillance data, the combination of the challenging effects is the biggest problem because typical counter measures are individual to one single effect. Christian Herrmann 0001, Chengchao Qu, Dieter Willersinn, Jürgen Beyerer |
AVSS | 4 |
| 2015 | Adaptive Contour Fitting for Pose-Invariant 3D Face Shape ReconstructionabstractDirect reconstruction of 3D face shape—solely based on a sparse set of 2D feature points localized by a facial landmark detector—offers an automatic, efficient and illumination-invariant alternative to the conventional analysis-by-synthesis 3D Morphable Model (3DMM) fitting. In this paper, we propose a novel algorithm that addresses the inconsistent correspondence of 2D and 3D landmarks at the facial contour due to head pose and localization ambiguity along the edge. To facilitate dynamic correspondence while fitting, a small subset of 3D vertices that serves as the contour candidates is annotated offline. During the fitting process, we employ the Levenberg-Marquardt Iterative Closest Point (LM-ICP) algorithm in combination with Distance Transform (DT) within the constrained domain, which allows for fast convergence and robust estimation of 3D face shape against pose variation. Superior evaluation results reported on ground truth 3D face scans over the state-of-the-art demonstrate the efficacy of the proposed method Chengchao Qu, Eduardo Monari, Tobias Schuchert, Jürgen Beyerer |
BMVC | 4 |
| 2015 | Correspondence between variational methods and Hidden Markov ModelsabstractThis paper establishes a duality between the calculus of variations, an increasingly common method for trajectory planning, and Hidden Markov Models (HMMs), a common probabilistic graphical model with applications in artificial intelligence and machine learning. This duality allows findings from each field to be applied to the other, namely providing an efficient and robust global optimization tool and machine learning algorithms for variational problems, and fast local solution methods for large state-space HMMs. Jens R. Ziehn, Miriam Ruf, Bodo Rosenhahn, Dieter Willersinn, Jürgen Beyerer, Heinrich Gotzig |
Intelligent Vehicles Symposium | 5 |
| 2014 | Cavlectometry: Towards Holistic Reconstruction of Large Mirror ObjectsabstractWe introduce a method based on the deflectometry principle for the reconstruction of specular objects exhibiting significant size and geometric complexity. A key feature of our approach is the deployment of an Automatic Virtual Environment (CAVE) as pattern generator. To unfold the full power of this experimental setup, an optical encoding scheme is developed which accounts for the distinctive topology of the CAVE. Furthermore, we devise an algorithm for detecting the object of interest in raw deflect metric images. The segmented foreground is used for single-view reconstruction, the background for estimation of the camera pose, necessary for calibrating the sensor system. Experiments suggest a significant gain of coverage in single measurements compared to previous methods. Jonathan Balzer, Daniel Acevedo Feliz, Stefano Soatto, Sebastian Höfer, Markus Hadwiger, Jürgen Beyerer |
3DV | 6 |
| 2014 | Maximizing face recognition performance for video data under time constraints by using a cascadeabstractThis paper presents a cascade of video face recognition methods which maximizes the recognition performance within a given time limit. Traditionally, choosing the appropriate recognition method for a specific person identification scenario is a difficult task. The maximization of the recognition performance under restricted time is seen as the main problem for increasing sizes of video data. To address this, we first evaluate a set of common face recognition methods and identify the ones which show the best recognition performance per time. Then, they are combined in the proposed cascade. An optimization strategy is presented, that maximizes the recognition performance of the cascade within a given time limit. Especially, this is no fixed limit, instead it can be assigned for each recognition task individually. A cross dataset evaluation on the Honda/UCSD and the Face in Action datasets shows the benefits of the proposed cascade. It allows situation specific configuration and the recognition performance per time is improved in comparison to the underlying face recognition methods. Christian Herrmann 0001, Jürgen Beyerer |
AVSS | 2 |
| 2014 | Fast, robust and automatic 3D face model reconstruction from videosabstractThis paper presents a fully automatic system that recovers 3D face models from sequences of facial images. Unlike most 3D Morphable Model (3DMM) fitting algorithms that simultaneously reconstruct the shape and texture from a single input image, our approach builds on a more efficient least squares method to directly estimate the 3D shape from sparse 2D landmarks, which are localized by face alignment algorithms. The inconsistency between self-occluded 2D and 3D feature positions caused by head pose is ad-dressed. A novel framework to enhance robustness across multiple frames selected based on their 2D landmarks combined with individual self-occlusion handling is proposed. Evaluation on groundtruth 3D scans shows superior shape and pose estimation over previous work. The whole system is also evaluated on an “in the wild” video dataset [12] and delivers personalized and realistic 3D face shape and texture models under less constrained conditions, which only takes seconds to process each video clip. Chengchao Qu, Eduardo Monari, Tobias Schuchert, Jürgen Beyerer |
AVSS | 4 |
| 2014 | Evaluation of object segmentation to improve moving vehicle detection in aerial videosabstractMoving objects play a key role for gaining scene understanding in aerial surveillance tasks. The detection of moving vehicles can be challenging due to high object distance, simultaneous object and camera motion, shadows, or weak contrast. In scenarios where vehicles are driving on busy urban streets, this is even more challenging due to possible merged detections. In this paper, a video processing chain is proposed for moving vehicle detection and segmentation. The fundament for detecting motion which is independent of the camera motion is tracking of local image features such as Harris corners. Independently moving features are clustered. Since motion clusters are prone to merge similarly moving objects, we evaluate various object segmentation approaches based on contour extraction, blob extraction, or machine learning to handle such effects. We propose to use a local sliding window approach with Integral Channel Features (ICF) and AdaBoost classifier. Michael Teutsch, Wolfgang Krüger, Jürgen Beyerer |
AVSS | 3 |
| 2014 | Comparing mouse and MAGIC pointing for moving target acquisitionabstractMoving target acquisition is a challenging and manually stressful task if performed using an all-manual, pointer-based interaction technique like mouse interaction, especially if targets are small, move fast, and are visible on screen only for a limited time. The MAGIC pointing interaction approach combines the precision of manual, pointer-based interaction with the speed and little manual stress of eye pointing. In this contribution, a pilot study with twelve participants on moving target acquisition is presented using an abstract experimental task derived from a video analysis scenario. Mouse input, conservative MAGIC pointing and MAGIC button are compared considering acquisition time, error rate, and user satisfaction. Although none of the participants had used MAGIC pointing before, eight participants voted for MAGIC button being their favorite technique; participants performed with only slightly higher mean acquisition time and error rate than with the familiar mouse input. Conservative MAGIC pointing was preferred by three participants; however, mean acquisition time and error rate were significantly worse than with mouse input. Jutta Hild, Dennis Gill, Jürgen Beyerer |
ETRA | 3 |
| 2014 | Realistic heatmap visualization for interactive analysis of 3D gaze dataabstractIn this paper, a novel approach for real-time heatmap generation and visualization of 3D gaze data is presented. By projecting the gaze into the scene and considering occlusions from the observer's view, to our knowledge, for the first time a correct visualization of the actual scene perception in 3D environments is provided. Based on a graphics-centric approach utilizing the graphics pipeline, shaders and several optimization techniques, heatmap rendering is fast enough for an interactive online and offline gaze analysis of thousands of gaze samples. Michael Maurus, Jan Hendrik Hammer, Jürgen Beyerer |
ETRA | 3 |
| 2014 | Optical filter selection for automatic visual inspectionabstractThe color of a material is one of the most frequently used features in automated visual inspection systems. While this is sufficient for many “easy” tasks, mixed and organic materials usually require more complex features. Spectral signatures, especially in the near infrared range, have been proven useful in many cases. However, hyperspectral imaging devices are still very costly and too slow to use them in practice. As a work-around, off-the-shelve cameras and optical filters are used to extract few characteristic features from the spectra. Often, these filters are selected by a human expert in a time consuming and error prone process; surprisingly few works are concerned with automatic selection of suitable filters. We approach this problem by stating filter selection as feature selection problem. In contrast to existing techniques that are mainly concerned with filter design, our approach explicitly selects the best out of a large set of given filters. Our method becomes most appealing for use in an industrial setting, when this selection represents (physically) available filters. We show the application of our technique by implementing six different selection strategies and applying each to two real-world sorting problems. Matthias Richter 0003, Jürgen Beyerer |
WACV | 2 |
| 2013 | PPRS: Production skills and their relation to product, process, and resourceabstractTo model increasingly adaptive production systems, skills are used to describe generic capabilities of the system components. In this paper, the authors extend the well-known division of production entities into product, process, and resource (PPR) with a skill definition. There are two main advantages for this approach: First, using PPR for the skill definition allows easy integration into existing models and tools. Second, there is a natural tendency to define very generic skills to capture all possible use cases. But at some point, skills have to be translated into precise instructions for execution. The model makes this dichotomy explicit and provides a common taxonomy for stakeholders concerned with skills on different abstraction levels. Julius Pfrommer, Miriam Schleipen, Jürgen Beyerer |
ETFA | 3 |
| 2013 | Adaptive real-time image smoothing using local binary patterns and Gaussian filtersabstractImage smoothing is widely used for enhancing the quality of single images or videos. There is a large amount of application areas such as machine vision, entertainment industry with smart TVs or consumer cameras, or surveillance and reconnaissance with different imaging sensors. In many cases it is not easy to find the trade-off between high smoothing quality and fast processing time. However, this is necessary for the mentioned applications as they are dependent on real-time computing. In this paper, we aim to find a good trade-off. Local texture is analyzed with Local Binary Patterns (LBPs) which are used to adapt the size of a Gaussian smoothing kernel for each pixel. Real-time requirements are met by the implementation on a Graphical Processing Unit (GPU). An image of 512 × 512 pixels is processed in 2.6 ms. Michael Teutsch, Patrick Trantelle, Jürgen Beyerer |
ICIP | 3 |
| 2013 | Locating user attention using eye tracking and EEG for spatio-temporal event selectionabstractIn expert video analysis, the selection of certain events in a continuous video stream is a frequently occurring operation, e.g., in surveillance applications. Due to the dynamic and rich visual input, the constantly high attention and the required hand-eye coordination for mouse interaction, this is a very demanding and exhausting task. Hence, relevant events might be missed. We propose to use eye tracking and electroencephalography (EEG) as additional input modalities for event selection. From eye tracking, we derive the spatial location of a perceived event and from patterns in the EEG signal we derive its temporal location within the video stream. This reduces the amount of the required active user input in the selection process, and thus has the potential to reduce the user's workload. In this paper, we describe the employed methods for the localization processes and introduce the developed scenario in which we investigate the feasibility of this approach. Finally, we present and discuss results on the accuracy and the speed of the method and investigate how the modalities interact. Felix Putze, Jutta Hild, Rainer Kärgel, Christian Herff, Alexander Redmann, Jürgen Beyerer, Tanja Schultz |
IUI | 6 |
| 2013 | Performance improvement of character recognition in industrial applications using prior knowledge for more reliable segmentation
Martin Grafmüller, Jürgen Beyerer |
Expert Syst. Appl. | 2 |
| 2012 | Defining dynamic Bayesian networks for probabilistic situation assessment
Yvonne Fischer, Jürgen Beyerer |
FUSION | 2 |
| 2012 | Bayesian active object recognition via Gaussian process regression
Marco F. Huber, Tobias Dencker, Masoud Roschani, Jürgen Beyerer |
FUSION | 4 |
| 2011 | Multiview specular stereo reconstruction of large mirror surfacesabstractIn deflectometry, the shape of mirror objects is recovered from distorted images of a calibrated scene. While remarkably high accuracies are achievable, state-of-the-art methods suffer from two distinct weaknesses: First, for mainly constructive reasons, these can only capture a few square centimeters of surface area at once. Second, reconstructions are ambiguous i.e. infinitely many surfaces lead to the same visual impression. We resolve both of these problems by introducing the first multiview specular stereo approach, which jointly evaluates a series of overlapping deflectometric images. Two publicly available benchmarks accompany this paper, enabling us to numerically demonstrate viability and practicability of our approach. Jonathan Balzer, Sebastian Höfer, Jürgen Beyerer |
CVPR | 3 |
| 2011 | Fusion of region and point-feature detections for measurement reconstruction in multi-target Kalman tracking
Michael Teutsch, Wolfgang Krüger, Jürgen Beyerer |
FUSION | 3 |
| 2011 | A comparison of motion planning algorithms for cooperative collision avoidance of multiple cognitive automobilesabstractAutomated cooperative collision avoidance of multiple vehicles is a promising approach to increase road safety in the future. This approach requires a real-time motion planner which computes cooperative maneuvers of multiple cognitive vehicles. As motion planning is a task of high computational complexity, computing times of the planner have to be traded off against solution quality. This contribution compares several cooperative motion planning algorithms with respect to these criteria. The considered algorithms are a tree search algorithm relying on precomputed lower bounds, the elastic band method, mixed-integer linear programming, and a priority-based approach. Success rates and computing times on various simulated scenarios are reported. Christian Frese, Jürgen Beyerer |
Intelligent Vehicles Symposium | 2 |
| 2010 | Lift-and-drop: crossing boundaries in a multi-display environment by AirliftabstractMany of the interactive environments surrounding us today consist of multiple mobile and/or stationary visual displays. However, interaction with such multi-display environments is still dominated by the personal computer paradigm - one user interacts with one single display at a time.In this paper first we present a new video-based input device called Airlift, which captures hands and fingertips independent from any display and therefore allows for consistent interaction across display boundaries. Second, we propose a system architecture for interaction spanning multiple displays. Third, we start to explore this new design space by proposing and evaluating a new interaction technique Lift-and-Drop for copying data from one display to another.According to the results of our study for the task considered the new technique is superior to other techniques based on traditional direct input devices which are limited to the surface of single displays like pen or touch. Thomas Bader, Astrid Heck, Jürgen Beyerer |
AVI | 3 |
| 2010 | Skill-based telemanipulation by means of intelligent robotsabstractIn order to enable robots to execute highly dynamic tasks in dangerous or remote environments, a semiautomatic teleoperation concept has been developed and will be presented in this paper. It relies on a modular software architecture, which allows intuitive control over the robot and compensates latency-based risks by using Augmented Reality techniques together with path prediction and collision avoidance to provide the remote user with visual feedback about the tasks and skills that will be executed. Based on this architecture different skills with high dynamics are integrated in the robot control, so that they can be executed autonomously without the delayed feedback of the user. The skill-based grasping by adherence of smooth or fragile objects during a remote controlled picking and placing task will be exemplary presented. Simon Notheis, Giulio Milighetti, Björn Hein, Heinz Wörn, Jürgen Beyerer |
IROS | 5 |
| 2009 | Fast Invariant Contour-Based Classification of Hand Symbols for HCI
Thomas Bader, René Räpple, Jürgen Beyerer |
CAIP | 3 |
| 2008 | Bayesian fusion of multivariate image to obtain depth information
Ioana Gheta, Michael Heizmann, Jürgen Beyerer |
FUSION | 3 |
| 2008 | Decreased complexity and increased problem specificity of Bayesian fusion by local approaches
Jennifer Sander, Jürgen Beyerer |
FUSION | 2 |
| 2008 | Shape from Specular Reflection and Optical Flow
Jan Lellmann, Jonathan Balzer, Andreas Rieder, Jürgen Beyerer |
Int. J. Comput. Vis. | 4 |
| 1998 | Is it useful to know a nuisance parameter?
Jürgen Beyerer |
Signal Process. | 1 |