Joachim Denzler

dblp:d/JoachimDenzler · DBLP profile ↗
← Back
112ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-3193-3300ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 70 · 7 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Realistic Face Reconstruction from Facial Embeddings via Diffusion Models
abstract
With the advancement of face recognition (FR) systems, privacy-preserving face recognition (PPFR) systems have gained popularity for their accurate recognition, enhanced facial privacy protection, and robustness to various attacks. However, there are limited studies to further verify privacy risks by reconstructing realistic high-resolution face images from embeddings of these systems, especially for PPFR. In this work, we propose the face embedding mapping (FEM), a general framework that explores Kolmogorov-Arnold Network (KAN) for conducting the embedding-to-face attack by leveraging pre-trained Identity-Preserving diffusion model against state-of-the-art (SOTA) FR and PPFR systems. Based on extensive experiments, we verify that reconstructed faces can be used for accessing other real-word FR systems. Besides, the proposed method shows the robustness in reconstructing faces from the partial and protected face embeddings. Moreover, FEM can be utilized as a tool for evaluating safety of FR and PPFR systems in terms of privacy leakage. All images used in this work are from public datasets.
Yong Li 0021, Joachim Denzler
AAAI3
2026 Locally Explaining Prediction Behavior via Gradual Interventions and Measuring Property Gradients
abstract
Deep learning models achieve high predictive performance but lack intrinsic interpretability, hindering our understanding of the learned prediction behavior. Existing local explainability methods focus on associations, neglecting the causal drivers of model predictions. Other approaches adopt a causal perspective but primarily provide global, model-level explanations. However, for specific inputs, it’s unclear whether globally identified factors apply locally. To address this limitation, we introduce a novel framework for local interventional explanations by leveraging recent advances in image-to-image editing models. Our approach performs gradual interventions on semantic properties to quantify the corresponding impact on a model’s predictions using a novel score, the expected property gradient magnitude. We demonstrate the effectiveness of our approach through an extensive empirical evaluation on a wide range of architectures and tasks. First, we validate it in a synthetic scenario and demonstrate its ability to locally identify biases. Afterward, we apply our approach to investigate medical skin lesion classifiers, analyze network training dynamics, and study a pre-trained CLIP model with real-life interventional data. Our results highlight the potential of interventional explanations on the property level to reveal new insights into the behavior of deep models.1
Niklas Penzel, Joachim Denzler
WACV2
2026 F-INR: Functional Tensor Decomposition for Implicit Neural Representations
abstract
Implicit Neural Representations (INRs) model signals as continuous, differentiable functions. However, monolithic INRs scale poorly with data dimensionality, leading to excessive training costs. We propose F-INR, a framework that addresses this limitation by factorizing a high-dimensional INR into a set of compact, axis-specific sub-networks based on functional tensor decomposition. These sub-networks learn low-dimensional functional components that are then combined via tensor operations. This factorization reduces computational complexity while additionally improving representational capacity. F-INR is both architecture-and decomposition-agnostic. It integrates with various existing INR backbones (e.g., SIREN, WIRE, FINER, Factor Fields) and tensor formats (e.g., CP, TT, Tucker), offering fine-grained control over the speed-accuracy trade-off via the tensor rank and mode. Our experiments show F-INR accelerates training by up to 20× and improves fidelity by over 6.0 dB PSNR compared to state-of-the-art INRs. We validate these gains on diverse tasks, including image representation, 3D geometry reconstruction, and neural radiance fields. We further show F-INR’s applicability to scientific computing by modeling complex physics simulations. Thus, F-INR provides a scalable, flexible, and efficient framework for high-dimensional signal modeling1.
Sai Karthikeya Vemuri, Tim Büchner, Joachim Denzler
WACV3
2026 Model utility and explainability in federated learning - A case study in healthcare using fundus oculi datasets
abstract
OBJECTIVE: Introduce a case study for Federated Learning (FL) in healthcare, addressing challenges posed by patient privacy and limited large-scale datasets. Our goal is to assess the features learned by FL methods in a simulated, diverse setting that emphasizes realistic data heterogeneity, and to analyze the learned representations for their medical relevance using both local and global explainability techniques. METHODS: Six fundus oculi datasets were combined to simulate a diverse federated learning environment, representing heterogeneous data conditions. We evaluated three established FL methods against centrally trained models, assessing both predictive performance and the learned representations. Specifically, explainability techniques were employed to examine the features learned by the models, and local explanations were evaluated against attention maps annotated by ophthalmologists. Robustness against common biases in fundus datasets was also assessed. RESULTS: Our study found improvements in model utility (up to 9.97%) with FL methods compared to isolated training. Analysis of learned representations revealed that federated models predominantly learn the vertical cup-to-disc ratio, a crucial feature for glaucoma diagnosis, and demonstrated robustness against common biases. High agreement was observed between local explanations and ophthalmologist-annotated attention maps. CONCLUSION: This study demonstrates the benefits of FL systems in a healthcare scenario, providing a case study for evaluating federated systems beyond idealized benchmarks. Our findings highlight the potential of FL to not only improve model utility in privacy-sensitive medical domains but also to learn medically relevant features instead of spurious correlations.
Niklas Penzel, Daniel Scheliga, Hannes Oppermann, Patrick Mäder, Jens Haueisen, Joachim Denzler, Marco Seeland
J. Biomed. Informatics6
2026 Scalable and expressive physics-informed neural networks via functional tensor decomposition
Sai Karthikeya Vemuri, Tim Büchner, Julia Niebling, Joachim Denzler
Pattern Recognit. Lett.4
2025 Adaptive Model Selection for Expanded Post Hoc Debiasing and Mitigating Varying Degrees of Spurious Correlations
Jan Blunk, Paul Bodesheim, Joachim Denzler
CAIP (2)3
2025 Electromyography-Informed Facial Expression Reconstruction for Physiological-Based Synthesis and Analysis
abstract
The relationship between muscle activity and resulting facial expressions is crucial for various fields, including psychology, medicine, and entertainment. The synchronous recording of facial mimicry and muscular activity via surface electromyography (sEMG) provides a unique window into these complex dynamics. Unfortunately, existing methods for facial analysis cannot handle electrode occlusion, rendering them ineffective. Even with occlusion-free reference images of the same person, variations in expression intensity and execution are unmatchable. Our electromyography-informed facial expression reconstruction (EIFER) approach is a novel method to restore faces under sEMG occlusion faithfully in an adversarial manner. We decouple facial geometry and visual appearance (e.g., skin texture, lighting, electrodes) by combining a 3D Morphable Model (3DMM) with neural unpaired image-to-image translation via reference recordings. Then, EIFER learns a bidirectional mapping between 3DMM expression parameters and muscle activity, establishing correspondence between the two domains. We validate the effectiveness of our approach through experiments on a dataset of synchronized sEMG recordings and facial mimicry, demonstrating faithful geometry and appearance reconstruction. Further, we synthesize expressions based on muscle activity and how observed expressions can predict dynamic muscle activity. Consequently, EIFER introduces a new paradigm for facial electromyography, which could be extended to other forms of multi-modal face recordings1.
Tim Büchner, Christoph Anders 0001, Orlando Guntinas-Lichius, Joachim Denzler
CVPR4
2025 Diffusion-based Identity-Preserving Facial Privacy Protection
abstract
The efficacy of facial recognition systems that utilize deep learning techniques has led to significant concerns over privacy, since they possess the capability to facilitate unauthorized monitoring of individuals in the digital realm. Current techniques for improving privacy are ineffective in producing "naturalistic" photographs that can safeguard facial features and fail to ensure privacy while maintaining an optimal user experience. We present an innovative text-agnostic method for protecting facial privacy. Our method depends on manipulating the sampling process of a pretrained diffusion model utilizing the guidance from a target image face together with the original image and face guidance in an adversarial manner to produce a protected face image. We preserve the original visual information from the input face image for identity preservation while extracting general embedding information from the target face image for soft facial attribute transfer. The output protected face image from our method has imperceptible facial changes with enhanced privacy protection against state-of-the-art (SOTA) face recognition (FR) systems. Our extensive studies have shown that the faces generated using our method have a higher level of black-box adaptability, resulting in an absolute improvement of 6.4% on CelebA-HQ compared to the current most effective SOTA facial privacy protection technique in the face verification task while maintaining high image fidelity.
Salaheldin Mohamed, Yong Li 0021, Joachim Denzler
ICASSP4
2025 Gradient Extrapolation for Debiased Representation Learning
Ihab Asaad, Maha Shadaydeh, Joachim Denzler
ICCV3
2025 CausalRivers - Scaling up benchmarking of causal discovery for real-world time-series
abstract
Causal discovery, or identifying causal relationships from observational data, is a notoriously challenging task, with numerous methods proposed to tackle it. Despite this, in-the-wild evaluation of these methods is still lacking, as works frequently rely on synthetic data evaluation and sparse real-world examples under critical theoretical assumptions. Real-world causal structures, however, are often complex, evolving over time, non-linear, and influenced by unobserved factors, making it hard to decide on a proper causal discovery strategy. To bridge this gap, we introduce CausalRivers, the largest in-the-wild causal discovery benchmarking kit for time-series data to date. CausalRivers features an extensive dataset on river discharge that covers the eastern German territory (666 measurement stations) and the state of Bavaria (494 measurement stations). It spans the years 2019 to 2023 with a 15-minute temporal resolution. Further, we provide additional data from a flood around the Elbe River, as an event with a pronounced distributional shift. Leveraging multiple sources of information and time-series meta-data, we constructed two distinct causal ground truth graphs (Bavaria and eastern Germany). These graphs can be sampled to generate thousands of subgraphs to benchmark causal discovery across diverse and challenging settings. To demonstrate the utility of CausalRivers, we evaluate several causal discovery approaches through a set of experiments to identify areas for improvement. CausalRivers has the potential to facilitate robust evaluations and comparisons of causal discovery methods. Besides this primary purpose, we also expect that this dataset will be relevant for connected areas of research, such as time-series forecasting and anomaly detection. Based on this, we hope to push benchmark-driven method development that fosters advanced techniques for causal discovery, as is the case for many other areas of machine learning.
Gideon Stein, Maha Shadaydeh, Jan Blunk, Niklas Penzel, Joachim Denzler
ICLR5
2025 FastCAV: Efficient Computation of Concept Activation Vectors for Explaining Deep Neural Networks
abstract
Concepts such as objects, patterns, and shapes are how humans understand the world. Building on this intuition, concept-based explainability methods aim to study representations learned by deep neural networks in relation to human-understandable concepts. Here, Concept Activation Vectors (CAVs) are an important tool and can identify whether a model learned a concept or not. However, the computational cost and time requirements of existing CAV computation pose a significant challenge, particularly in large-scale, high-dimensional architectures. To address this limitation, we introduce FastCAV, a novel approach that accelerates the extraction of CAVs by up to 63.6× (on average 46.4×). We provide a theoretical foundation for our approach and give concrete assumptions under which it is equivalent to established SVM-based methods. Our empirical results demonstrate that CAVs calculated with FastCAV maintain similar performance while being more efficient and stable. In downstream applications, i.e., concept-based explanation methods, we show that FastCAV can act as a replacement leading to equivalent insights. Hence, our approach enables previously infeasible investigations of deep models, which we demonstrate by tracking the evolution of concepts during model training.
Laines Schmalwasser, Niklas Penzel, Joachim Denzler, Julia Niebling
ICML3
2025 Distance-informed Neural Processes
abstract
We propose the Distance-informed Neural Process (DNP), a novel variant of Neural Processes that improves uncertainty estimation by combining global and distance-aware local latent structures. Standard Neural Processes (NPs) often rely on a global latent variable and struggle with uncertainty calibration and capturing local data dependencies. DNP addresses these limitations by introducing a global latent variable to model task-level variations and a local latent variable to capture input similarity within a distance-preserving latent space. This is achieved through bi-Lipschitz regularization, which bounds distortions in input relationships and encourages the preservation of relative distances in the latent space. This modeling approach allows DNP to produce better-calibrated uncertainty estimates and more effectively distinguish in- from out-of-distribution data. Empirical results demonstrate that DNP achieves strong predictive performance and improved uncertainty calibration across regression and classification tasks.
Aishwarya Venkataramanan, Joachim Denzler
NeurIPS2
2025 Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models
abstract
Vision-Language Models (VLMs) learn joint representations by mapping images and text into a shared latent space. However, recent research highlights that deterministic embeddings from standard VLMs often struggle to capture the uncertainties arising from the ambiguities in visual and textual descriptions and the multiple possible correspondences between images and texts. Existing approaches tackle this by learning probabilistic embeddings during VLM training, which demands large datasets and does not leverage the powerful representations already learned by large-scale VLMs like CLIP. In this paper, we propose GroVE, a post-hoc approach to obtaining probabilistic embeddings from frozen VLMs. GroVE builds on Gaussian Process Latent Variable Model (GPLVM) to learn a shared low-dimensional latent space where image and text inputs are mapped to a unified representation, optimized through single-modal embedding reconstruction and cross-modal alignment objectives. Once trained, the Gaussian Process model generates uncertainty-aware probabilistic embeddings. Evaluation shows that GroVE achieves state-of-the-art uncertainty calibration across multiple downstream tasks, including cross-modal retrieval, visual question answering, and active learning.
Aishwarya Venkataramanan, Paul Bodesheim, Joachim Denzler
UAI3
2025 Simplified Concrete Dropout - Improving the Generation of Attribution Masks for Fine-grained Classification
abstract
Abstract In fine-grained classification, which is classifying images into subcategories within a common broader category, it is crucial to have precise visual explanations of the classification model’s decision. While commonly used attention- or gradient-based methods deliver either too coarse or too noisy explanations unsuitable for highlighting subtle visual differences reliably, perturbation-based methods can precisely locate pixels causally responsible for the predicted category. The fill-in of the dropout (FIDO) algorithm is one of those methods, which utilizes concrete dropout (CD) to sample a set of attribution masks and updates the sampling parameters based on the output of the classification model. In this paper, we present a solution against the high variance in the gradient estimates, a known problem of the FIDO algorithm that has been mitigated until now by large mini-batch updates of the sampling parameters. First, our solution allows for estimating the parameters with smaller mini-batch sizes without losing the quality of the estimates but with a reduced computational effort. Next, our method produces finer and more coherent attribution masks. Finally, we use the resulting attribution masks to improve the classification performance on three fine-grained datasets without additional fine-tuning steps and achieve results that are otherwise only achieved if ground truth bounding boxes are used.
Dimitri Korsch, Maha Shadaydeh, Joachim Denzler
Int. J. Comput. Vis.3
2025 Assessing 3D volumetric asymmetry in facial palsy patients via advanced multi-view landmarks and radial curves
abstract
Abstract The research on facial palsy, a unilateral palsy of the facial nerve, is a complex field with many different causes and symptoms. Even modern approaches to evaluate the facial palsy state rely mainly on stills and 2D videos of the face and rarely on dynamic 3D information. Many of these analysis and visualization methods require manual intervention, which is time-consuming and error-prone. Moreover, they often depend on alignment algorithms or Euclidean measurements and consider only static facial expressions. Volumetric changes by muscle movement are essential for facial palsy analysis but require manual extraction. We propose to extract an estimated unilateral volumetric description for dynamic expressions from 3D scans. Accurate landmark positioning is required for processing the unstructured facial scans. In our case, it is attained via a multi-view method compatible with any existing 2D predictors. We analyze prediction stability and robustness against head rotation during video sequences. Further, we investigate volume changes in static and dynamic facial expressions for 34 patients with unilateral facial palsy and visualize volumetric disparities on the face surface. In a case study, we observe a decrease in the volumetric difference between the face sides during happy expressions at the beginning (13.8 ± 10.0 $$\hbox {mm}^{3}$$ mm 3 ) and end (12.8 ± 10.3 $$\hbox {mm}^{3}$$ mm 3 ) of a ten-day biofeedback therapy. The neutral face kept a consistent volume range of 11.8 $$-$$ - 12.1 $$\hbox {mm}^3$$ mm 3 . The reduced volumetric difference after therapy indicates less facial asymmetry during movement, which can be used to monitor and guide treatment decisions. Our approach minimizes human intervention, simplifying the clinical routine and interaction with 3D scans to provide a more comprehensive analysis of facial palsy.
Tim Büchner, Sven Sickert, Gerd Fabian Volk, Orlando Guntinas-Lichius, Joachim Denzler
Mach. Vis. Appl.5
2024 Robust Skin Color Driven Privacy-Preserving Face Recognition Via Function Secret Sharing
abstract
In this work, we leverage the pure skin color patch from the face image as the additional information to train an auxiliary skin color feature extractor and face recognition model in parallel to improve performance of state-of-the-art (SOTA) privacy-preserving face recognition (PPFR) systems. Our solution is robust against black-box attacking and well-established generative adversarial network (GAN) based image restoration. We analyze the potential risk in previous work, where the proposed cosine similarity computation might directly leak the protected precomputed embedding stored on the server side. We propose a Function Secret Sharing (FSS) based face embedding comparison protocol without any intermediate result leakage. In addition, we show in experiments that the proposed protocol is more efficient compared to the Secret Sharing (SS) based protocol.
Yufan Jiang, Yong Li 0021, Ricardo Mendes, Joachim Denzler
ICIP5
2024 Exploiting Text-Image Latent Spaces for the Description of Visual Concepts
Laines Schmalwasser, Jakob Gawlikowski, Joachim Denzler, Julia Niebling
ICPR (33)3
2024 Functional Tensor Decompositions for Physics-Informed Neural Networks
Sai Karthikeya Vemuri, Tim Büchner, Julia Niebling, Joachim Denzler
ICPR (25)4
2024 Data-Driven Prediction Of Large Infrastructure Movements Through Persistent Scatterer Time Series Modeling
abstract
Deformation monitoring is a crucial task for dam operators, particularly given the rise in extreme weather events associated with climate change. Further, quantifying the expected deformations of a dam is a central part of this endeavor. Current methods rely on in situ data (i.e., water level and temperature) to predict the expected deformations of a dam (typically represented by plumb or trigonometric measurements). However, not all dams are equipped with extensive measurement techniques, resulting in infrequent monitoring. Persistent Scatterer Interferometry (PSI) can overcome this limitation, enabling an alternative monitoring scheme for such infrastructures. This study introduces a novel monitoring approach to quantify expected deformations of gravity dams in Germany by integrating the PSI technique with in situ data. Further, it proposes a methodology to find proper statistical representations in a data-driven manner, which extends established statistical approaches. The approach demonstrates plausible deformation patterns as well as accurate predictions for validation data (mean absolute error=1.81 mm), confirming the benefits of the proposed method.
Gideon Stein, Jonas Ziemer, Carolin Wicker, Jannik Jänichen, Gabriele Demisch, Daniel Klöpper, Katja Last, Joachim Denzler, Christiane Schmullius, Maha Shadaydeh, Clémence Dubois
IGARSS8
2024 Unraveling Anomalies in Time: Unsupervised Discovery and Isolation of Anomalous Behavior in Bio-Regenerative Life Support System Telemetry
Ferdinand Rewicki, Jakob Gawlikowski, Julia Niebling, Joachim Denzler
ECML/PKDD (9)4
2023 Improved Obstructed Facial Feature Reconstruction for Emotion Recognition with Minimal Change CycleGANs
Tim Büchner, Orlando Guntinas-Lichius, Joachim Denzler
ACIVS3
2022 Facing Asymmetry - Uncovering the Causal Link Between Facial Symmetry and Expression Classifiers Using Synthetic Interventions
Tim Büchner, Niklas Penzel, Orlando Guntinas-Lichius, Joachim Denzler
ACCV (4)4
2022 Occlusion-Robustness of Convolutional Neural Networks via Inverted Cutout
abstract
Convolutional Neural Networks (CNNs) are able to reliably classify objects in images if they are clearly visible and only slightly affected by small occlusions. However, heavy occlusions can strongly deteriorate the performance of CNNs, which is critical for tasks where correct identification is paramount. For many real-world applications, images are taken in unconstrained environments under suboptimal conditions, where occluded objects are inevitable. We propose a novel data augmentation method called Inverted Cutout, which can be used for training a CNN by showing only small patches of the images. Together with this augmentation method, we present several ways of making the network robust against occlusion. On the one hand, we utilize a spatial aggregation module without modifying the base network and on the other hand, we achieve occlusion-robustness with appropriate fine-tuning in conjunction with Inverted Cutout. In our experiments, we compare two different aggregation modules and two loss functions on the Occluded-Vehicles and Occluded-COCO-Vehicles datasets, showing that our approach outperforms existing state-of-the-art methods for object categorization under varying levels of occlusion.
Matthias Körschens, Paul Bodesheim, Joachim Denzler
ICPR3
2022 Generative adversarial networks for biomedical time series forecasting and imputation
abstract
In the present systematic review we identified and summarised current research activities in the field of time series forecasting and imputation with the help of generative adversarial networks (GANs). We differentiate between imputation which describes the filling of missing values at intermediate steps and forecasting defining the prediction of future values. Especially the utilisation of such methods in the biomedical domain was to be investigated. To this end, 1057 publications were identified with the help of PubMed, Web of Science and Scopus. All studies that describe the use of GANs for the imputation/forecasting of time series were included irrespective of the application domain. Finally, 33 records were identified as eligible and grouped according to the topologies, losses, inputs and outputs of the presented GANs. In combination with a summary of all described application domains, this grouping served as a basis for analysing the peculiarities of the method in the biomedical context. Due to the broad spectrum of biomedical research, nearly all recognised methodologies are also applied in this domain. We could not identify any approach that proved itself superior in the biomedical area. Although GANs were initially designed to work in the image domain, many publications show that they are capable of imputing/forecasting non-visual time series.
Sven Festag, Joachim Denzler, Cord Spreckelsen
J. Biomed. Informatics2
2021 Causal Inference in Non-linear Time-series using Deep Networks and Knockoff Counterfactuals
abstract
Estimating causal relations is vital in understanding the complex interactions in multivariate time series. Non-linear coupling of variables is one of the major challenges in accurate estimation of cause-effect relations. In this paper, we propose to use deep autoregressive networks (DeepAR) in tandem with counterfactual analysis to infer nonlinear causal relations in multivariate time series. We extend the concept of Granger causality using probabilistic forecasting with DeepAR. Since deep networks can neither handle missing input nor out-of-distribution intervention, we propose to use the Knockoffs framework (Barber and Candès, 2015) for generating intervention variables and consequently counterfactual probabilistic forecasting. Knockoff samples are independent of their output given the observed variables and exchangeable with their counterpart variables without changing the underlying distribution of the data. We test our method on synthetic as well as real-world time series datasets. Overall our method outperforms the widely used vector autoregressive Granger causality and PCMCI in detecting nonlinear causal dependency in multivariate time series.
Maha Shadaydeh, Joachim Denzler
ICMLA3
2021 Anomaly Attribution of Multivariate Time Series using Counterfactual Reasoning
abstract
There are numerous methods for detecting anomalies in time series, but that is only the first step to understanding them. We strive to exceed this by explaining those anomalies. Thus we develop a novel attribution scheme for multivariate time series relying on counterfactual reasoning. We aim to answer the counterfactual question of would the anomalous event have occurred if the subset of the involved variables had been more similarly distributed to the data outside of the anomalous interval. Specifically, we detect anomalous intervals using the Maximally Divergent Interval (MDI) algorithm, replace a subset of variables with their in-distribution values within the detected interval and observe if the interval has become less anomalous, by re-scoring it with MDI. We evaluate our method on multivariate temporal and spatio-temporal data and confirm the accuracy of our anomaly attribution of multiple well-understood extreme climate events such as heatwaves and hurricanes.
Violeta Teodora Trifunov, Maha Shadaydeh, Björn Barz, Joachim Denzler
ICMLA4
2020 Determining the Relevance of Features for Deep Neural Networks
Christian Reimers, Jakob Runge, Joachim Denzler
ECCV (26)3
2020 Making Every Label Count: Handling Semantic Imprecision by Integrating Domain Knowledge
abstract
Noisy data, crawled from the web or supplied by volunteers such as Mechanical Turkers or citizen scientists, is considered an alternative to professionally labeled data. There has been research focused on mitigating the effects of label noise. It is typically modeled as inaccuracy, where the correct label is replaced by an incorrect label from the same set. We consider an additional dimension of label noise: imprecision. For example, a non-breeding snow bunting is labeled as a bird. This label is correct, but not as precise as the task requires. Standard softmax classifiers cannot learn from such a weak label because they consider all classes mutually exclusive, which non-breeding snow bunting and bird are not. We propose CHILLAX (Class Hierarchies for Imprecise Label Learning and Annotation eXtrapolation), a method based on hierarchical classification, to fully utilize labels of any precision. Experiments on noisy variants of NABirds and ILSVRC2012 show that our method outperforms strong baselines by as much as 16.4 percentage points, and the current state of the art by up to 3.9 percentage points.
Clemens-Alexander Brust, Björn Barz, Joachim Denzler
ICPR3
2020 Single-Shot 3D Detection of Vehicles from Monocular RGB Images via Geometrically Constrained Keypoints in Real-Time
abstract
In this paper we propose a novel 3D single-shot object detection method for detecting vehicles in monocular RGB images. Our approach lifts 2D detections to 3D space by predicting additional regression and classification parameters and hence keeping the runtime close to pure 2D object detection. The additional parameters are transformed to 3D bounding box keypoints within the network under geometric constraints. Our proposed method features a full 3D description including all three angles of rotation without supervision by any labeled ground truth data for the object's orientation, as it focuses on certain keypoints within the image plane. While our approach can be combined with any modern object detection framework with only little computational overhead, we exemplify the extension of SSD for the prediction of 3D bounding boxes. We test our approach on different datasets for autonomous driving and evaluate it using the challenging KITTI 3D Object Detection as well as the novel nuScenes Object Detection benchmarks. While we achieve competitive results on both benchmarks we outperform current state-of-the-art methods in terms of speed with more than 20 FPS for all tested datasets and image resolutions.
Nils Gählert, Jun-Jun Wan, Nicolas Jourdan 0001, Jan Finkbeiner, Uwe Franke, Joachim Denzler
IV6
2020 Deep Learning on Small Datasets without Pre-Training using Cosine Loss
abstract
Two things seem to be indisputable in the contemporary deep learning discourse: 1. The categorical cross-entropy loss after softmax activation is the method of choice for classification. 2. Training a CNN classifier from scratch on small datasets does not work well.In contrast to this, we show that the cosine loss function provides substantially better performance than crossentropy on datasets with only a handful of samples per class. For example, the accuracy achieved on the CUB- 200-2011 dataset without pre-training is by 30% higher than with the cross-entropy loss. Further experiments on other popular datasets confirm our findings. Moreover, we demonstrate that integrating prior knowledge in the form of class hierarchies is straightforward with the cosine loss and improves classification performance further.
Björn Barz, Joachim Denzler
WACV2
2020 The Whole Is More Than Its Parts? From Explicit to Implicit Pose Normalization
abstract
Fine-grained classification describes the automated recognition of visually similar object categories like birds species. Previous works were usually based on explicit pose normalization, i.e., the detection and description of object parts. However, recent models based on a final global average or bilinear pooling have achieved a comparable accuracy without this concept. In this paper, we analyze the advantages of these approaches over generic CNNs and explicit pose normalization approaches. We also show how they can achieve an implicit normalization of the object pose. A novel visualization technique called activation flow is introduced to investigate limitations in pose handling in traditional CNNs like AlexNet and VGG. Afterward, we present and compare the explicit pose normalization approach neural activation constellations and a generalized framework for the final global average and bilinear pooling called α-pooling. We observe that the latter often achieves a higher accuracy improving common CNN models by up to 22.9 percent, but lacks the interpretability of the explicit approaches. We present a visualization approach for understanding and analyzing predictions of the model to address this issue. Furthermore, we show that our approaches for fine-grained recognition are beneficial for other fields like action recognition.
Marcel Simon, Erik Rodner, Trevor Darrell, Joachim Denzler
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Edge-Convolution Point Net for Semantic Segmentation of Large-Scale Point Clouds
abstract
In this paper, we propose a deep learning-based framework which can manage large-scale point clouds of outdoor scenes with high spatial resolution. For large and high-resolution outdoor scenes, point-wise classification approaches are often an intractable problem. Analogous to Object-Based Image Analysis (OBIA), our approach segments the scene by grouping similar points together to generate meaningful objects. Later, our net classifies segments instead of individual points using an architecture inspired by PointNet, which applies Edge convolutions. This approach is trained using both visual and geometrical information. Experiments show the potential of this task even for small training sets. Furthermore, we can show competitive performance on a Large-scale Point Cloud Classification Benchmark.
Jhonatan Contreras, Joachim Denzler
IGARSS2
2019 Registration of High Resolution Sar and Optical Satellite Imagery Using Fully Convolutional Networks
abstract
Multi-modal image registration is a crucial step when fusing images which show different physical/chemical properties of an object. Depending on the compared modalities and the used registration metric, this process exhibits varying reliability. We propose a deep metric based on a fully convo-lutional neural network (FCN). It is trained from scratch on SAR-optical image pairs to predict whether certain image areas are aligned or not. Tests on the affine registration of SAR and optical images showing suburban areas verify an enormous improvement of the registration accuracy in comparison to registration metrics that are based on mutual information (MI).
Stefan Hoffmann 0007, Clemens-Alexander Brust, Maha Shadaydeh, Joachim Denzler
IGARSS4
2019 Beyond Bounding Boxes: Using Bounding Shapes for Real-Time 3D Vehicle Detection from Monocular RGB Images
abstract
The representation of objects as 2D bounding boxes in monocular RGB images limits the faculty of current computer vision systems to 2D object detection. It fails to provide crucial information such as the orientation of other vehicles, which is vital for autonomous driving. At the same time, real-time performance is essential to qualify an approach for deployment in a productive environment. In order to tackle this problem, we present an approach that predicts several key points selected from a virtual 3D bounding box around a vehicle instead of a pure 2D bounding box. These key points can be interpreted as a bounding shape. With this novel representation we can calculate the actual 3D bounding box of the corresponding object. Thanks to the straightforward implementation of bounding shape in any current state-of-the-art 2D object detector both for singleshot frameworks like YOLO or SSD as well as for two-stage detectors like Faster-RCNN with a minimum of computational overhead, it is able to be run in real-time while providing additional useful information for vehicle detection. We exemplify the extension of SSD to Bounding Shape SSD ( BS3D) and evaluate our approach using the challenging KITTI as well as the novel VIPER dataset.
Nils Gählert, Jun-Jun Wan, Michael Weber 0009, Johann Marius Zöllner, Uwe Franke, Joachim Denzler
IV6
2019 Hierarchy-Based Image Embeddings for Semantic Image Retrieval
abstract
Deep neural networks trained for classification have been found to learn powerful image representations, which are also often used for other tasks such as comparing images w.r.t. their visual similarity. However, visual similarity does not imply semantic similarity. In order to learn semantically discriminative features, we propose to map images onto class embeddings whose pair-wise dot products correspond to a measure of semantic similarity between classes. Such an embedding does not only improve image retrieval results, but could also facilitate integrating semantics for other tasks, e.g., novelty detection or few-shot learning. We introduce a deterministic algorithm for computing the class centroids directly based on prior world-knowledge encoded in a hierarchy of classes such as WordNet. Experiments on CIFAR-100, NABirds, and ImageNet show that our learned semantic image embeddings improve the semantic consistency of image retrieval results by a large margin.
Björn Barz, Joachim Denzler
WACV2
2019 Detecting Regions of Maximal Divergence for Spatio-Temporal Anomaly Detection
abstract
Automatic detection of anomalies in space- and time-varying measurements is an important tool in several fields, e.g., fraud detection, climate analysis, or healthcare monitoring. We present an algorithm for detecting anomalous regions in multivariate spatio-temporal time-series, which allows for spotting the interesting parts in large amounts of data, including video and text data. In opposition to existing techniques for detecting isolated anomalous data points, we propose the "Maximally Divergent Intervals" (MDI) framework for unsupervised detection of coherent spatial regions and time intervals characterized by a high Kullback-Leibler divergence compared with all other data given. In this regard, we define an unbiased Kullback-Leibler divergence that allows for ranking regions of different size and show how to enable the algorithm to run on large-scale data sets in reasonable time using an interval proposal technique. Experiments on both synthetic and real data from various domains, such as climate analysis, video surveillance, and text forensics, demonstrate that our method is widely applicable and a valuable tool for finding interesting events in different types of data.
Björn Barz, Erik Rodner, Yanira Guanche, Joachim Denzler
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Active Learning for Regression Tasks with Expected Model Output Changes
Christoph Käding, Erik Rodner, Alexander Freytag, Oliver Mothes, Björn Barz, Joachim Denzler
BMVC6
2018 Predicting Landscapes as Seen from Space from Environmental Conditions
abstract
Satellite images are information rich snapshots of ecosystems and landscapes. In consequence, the features in the images strongly depend on the environmental conditions. Such dependency between climate and landscapes has been regarded since the beginning of earth sciences; however, it has never been taken as literally as in the present study. We adapted a deep learning generative model as a first demonstration of the potential behind deep learning for spatial pattern generation in geoscience. The purpose is to build a conditional Generative Adversarial Network (cGAN) useful to establish the relationship between two loosely linked set of variables that show multitude of complex spatial features such as climate conditions to aerial image. We trained a custom cGAN to generate Sentinel-2 multispectral imagery given a set of climatic and terrain predictors. Results show that the generated imagery shares many characteristics with the real one. In some cases, the quality of the generated imagery is high enough to deceive humans. We envision that such use of deep learning for geoscience could become an important tool to test the effects of climate on landscapes and ecosystems.
Christian Requena-Mesa, Markus Reichstein, Miguel D. Mahecha, Basil Kraft, Joachim Denzler
IGARSS5
2018 MB-Net: MergeBoxes for Real-Time 3D Vehicles Detection
abstract
High performance vehicle detection and pose esti- mation in RGB images is essential for driver assistance systems as well as for autonomous vehicles. Classical 2D box-based detection schemes allow roughly estimating the position of other vehicles, but not their orientation relative to the ego-vehicle. Recent approaches use 3D models to derive the pose of other vehicles from single monocular images but do not reach real- time performance. In this paper we present an approach that achieves competitive performance on the challenging KITTI Object Detection and orientation Estimation benchmark while being the fastest approach with over 40 FPS. The key is a novel representation named MergeBox whose parameters can be estimated extremely efficiently. We extend SSD-a current fast state-of-the-art 2D box object detector- with this representation to our MB-Net. In contrast to all other current state-of-the-art methods we do not require explicit information on the object orientation for training our model. This reduces label costs significantly, a further advantage for practical applications that require labeling of databases that are much bigger than those used for research.
Nils Gählert, Marina Mayer, Lukas Schneider, Uwe Franke, Joachim Denzler
Intelligent Vehicles Symposium5
2017 A Feedback Estimation Approach for Therapeutic Facial Training
abstract
Neuromuscular retraining is an important part of facial paralysis rehabilitation. To date, few publications have addressed the development of automated systems that support facial training. Current approaches require external devices attached to the patient's face, lack quantitative feedback, and are constrained to one or two facial training exercises. We propose an automated camera-based training system that provides global and local feedback for 12 different facial training exercises. Based on extracted 3D facial features, the patient's performance is evaluated and quantitative feedback is derived. The description of the feedback estimation is supplemented by a detailed experimental evaluation of the 3D feature extraction.
Cornelia Dittmar, Joachim Denzler, Horst-Michael Groß
FG2
2017 Generalized Orderless Pooling Performs Implicit Salient Matching
abstract
Most recent CNN architectures use average pooling as a final feature encoding step. In the field of fine-grained recognition, however, recent global representations like bilinear pooling offer improved performance. In this paper, we generalize average and bilinear pooling to “α-pooling”, allowing for learning the pooling strategy during training. In addition, we present a novel way to visualize decisions made by these approaches. We identify parts of training images having the highest influence on the prediction of a given test image. This allows for justifying decisions to users and also for analyzing the influence of semantic parts. For example, we can show that the higher capacity VGG16 model focuses much more on the bird's head than, e.g., the lower-capacity VGG-M model when recognizing fine-grained bird categories. Both contributions allow us to analyze the difference when moving between average and bilinear pooling. In addition, experiments show that our generalized approach can outperform both across a variety of standard datasets.
Marcel Simon, Yang Gao 0029, Trevor Darrell, Joachim Denzler, Erik Rodner
ICCV4
2017 Towards unconstrained content recognition of additional traffic signs
abstract
The task of traffic sign recognition is often considered to be solved after almost perfect results have been achieved on some public benchmarks. Yet, the closely related recognition of additional traffic signs is still lacking a solution. Following up on our earlier work on detecting additional traffic signs given a main sign detection [1], we here propose a complete pipeline for recognizing the content of additional signs, including text recognition by optical character recognition (OCR). We assume a given additional sign detection, first classify its layout, then determine content bounding boxes by regression, followed by a multi-class classification step or, if necessary, OCR by applying a text sequence classifier. We evaluate the individual stages of our proposed pipeline and the complete system on a database of German additional signs and show that it can successfully recognize about 80% of the signs correctly, even under very difficult conditions and despite low input resolutions at runtimes well below 12ms per sign.
Thomas Wenzel, Steffen Brüggert, Joachim Denzler
Intelligent Vehicles Symposium3
2017 From corners to rectangles - Directional road sign detection using learned corner representations
abstract
In this work we adopt a novel approach for the detection of rectangular directional road signs in single frames captured from a moving car. These signs exhibit wide variations in sizes and aspect ratios and may contain arbitrary information, thus making their detection a challenging task with applications in traffic sign recognition systems and vision-based localization. Our proposed approach was originally presented for additional traffic sign detection in small image regions and is generalized to full image frames in this work. Sign corner areas are detected by four ACF-detectors (Aggregated Channel Features) on a single scale. The resulting corner detections are subsequently used to generate quadrangle hypotheses, followed by an aggressive pruning strategy. A comparative evaluation on a database of 1500 German road signs shows that our proposed detector outperforms other methods significantly at close to real-time runtimes and yields thrice the very low error-rate of the recent MS-CNN framework while being two orders of magnitude faster.
Thomas Wenzel, Ta-Wei Chou, Steffen Brüggert, Joachim Denzler
Intelligent Vehicles Symposium4
2017 Large-Scale Gaussian Process Inference with Generalized Histogram Intersection Kernels for Visual Recognition Tasks
Erik Rodner, Alexander Freytag, Paul Bodesheim, Björn Fröhlich, Joachim Denzler
Int. J. Comput. Vis.5
2017 Multi-marker tracking for large-scale X-ray stereo video data
Marcel Simon, Joachim Denzler
Signal Process. Image Commun.4
2016 Vegetation Segmentation in Cornfield Images Using Bag of Words
Yerania Campos, Erik Rodner, Joachim Denzler, Juan Humberto Sossa Azuela, Gonzalo Pajares
ACIVS3
2016 Impatient DNNs - Deep Neural Networks with Dynamic Time Budgets
Manuel Amthor, Erik Rodner, Joachim Denzler
BMVC3
2016 Fine-grained Recognition in the Noisy Wild: Sensitivity Analysis of Convolutional Neural Networks Approaches
Erik Rodner, Marcel Simon, Robert B. Fisher, Joachim Denzler
BMVC4
2016 Watch, Ask, Learn, and Improve: a lifelong learning cycle for visual recognition
Christoph Käding, Erik Rodner, Alexander Freytag, Joachim Denzler
ESANN4
2016 Additional traffic sign detection using learned corner representations
abstract
The detection of traffic signs and recognizing their meanings is crucial for applications such as online detection in automated driving or automated map data updates. Despite all progress in this field detecting and recognizing additional traffic signs, which may invalidate main traffic signs, has been widely disregarded in the scientific community. As a continuation of our earlier work we present a novel high-performing additional sign detector here, which outperforms our recently published state-of-the-art results significantly. Our approach relies on learning corner area representations using Aggregated Channel Features (ACF). Subsequently, a quadrangle generation and filtering strategy is applied, thus effectively dealing with the large aspect ratio variations of additional signs. It yields very high detection rates on a challenging dataset of high-resolution images captured with a windshield-mounted smartphone, and offers very precise localization while maintaining real-time capability. More than 95% of the additional traffic signs are detected successfully with full content detection at a false positive rate well below 0.1 per main sign, thus contributing a small step towards enabling automated driving.
Thomas Wenzel, Steffen Brüggert, Joachim Denzler
Intelligent Vehicles Symposium3
2015 Active learning and discovery of object categories in the presence of unnameable instances
abstract
Current visual recognition algorithms are “hungry” for data but massive annotation is extremely costly. Therefore, active learning algorithms are required that reduce labeling efforts to a minimum by selecting examples that are most valuable for labeling. In active learning, all categories occurring in collected data are usually assumed to be known in advance and experts should be able to label every requested instance. But do these assumptions really hold in practice? Could you name all categories in every image?
Christoph Käding, Alexander Freytag, Erik Rodner, Paul Bodesheim, Joachim Denzler
CVPR5
2015 Local Novelty Detection in Multi-class Recognition Problems
abstract
In this paper, we propose using local learning for multiclass novelty detection, a framework that we call local novelty detection. Estimating the novelty of a new sample is an extremely challenging task due to the large variability of known object categories. The features used to judge on the novelty are often very specific for the object in the image and therefore we argue that individual novelty models for each test sample are important. Similar to human experts, it seems intuitive to first look for the most related images thus filtering out unrelated data. Afterwards, the system focuses on discovering similarities and differences to those images only. Therefore, we claim that it is beneficial to solely consider training images most similar to a test sample when deciding about its novelty. Following the principle of local learning, for each test sample a local novelty detection model is learned and evaluated. Our local novelty score turns out to be a valuable indicator for deciding whether the sample belongs to a known category from the training set or to a new, unseen one. With our local novelty detection approach, we achieve state-of-the-art performance in multi-class novelty detection on two popular visual object recognition datasets, Caltech-256 and Image Net. We further show that our framework: (i) can be successfully applied to unknown face detection using the Labeled-Faces-in-the-Wild dataset and (ii) outperforms recent work on attribute-based unfamiliar class detection in fine-grained recognition of bird species on the challenging CUB-200-2011 dataset.
Paul Bodesheim, Alexander Freytag, Erik Rodner, Joachim Denzler
WACV4
2014 Part Detector Discovery in Deep Convolutional Neural Networks
Marcel Simon, Erik Rodner, Joachim Denzler
ACCV (2)3
2014 Nonparametric Part Transfer for Fine-Grained Recognition
abstract
In the following paper, we present an approach for fine-grained recognition based on a new part detection method. In particular, we propose a nonparametric label transfer technique which transfers part constellations from objects with similar global shapes. The possibility for transferring part annotations to unseen images allows for coping with a high degree of pose and view variations in scenarios where traditional detection models (such as deformable part models) fail. Our approach is especially valuable for fine-grained recognition scenarios where intraclass variations are extremely high, and precisely localized features need to be extracted. Furthermore, we show the importance of carefully designed visual extraction strategies, such as combination of complementary feature types and iterative image segmentation, and the resulting impact on the recognition performance. In experiments, our simple yet powerful approach achieves 35.9% and 57.8% accuracy on the CUB-2010 and 2011 bird datasets, which is the current best performance for these benchmarks.
Christoph Göring, Erik Rodner, Alexander Freytag, Joachim Denzler
CVPR4
2014 Instance-Weighted Transfer Learning of Active Appearance Models
abstract
There has been a lot of work on face modeling, analysis, and landmark detection, with Active Appearance Models being one of the most successful techniques. A major drawback of these models is the large number of detailed annotated training examples needed for learning. Therefore, we present a transfer learning method that is able to learn from related training data using an instance-weighted transfer technique. Our method is derived using a generalization of importance sampling and in contrast to previous work we explicitly try to tackle the transfer already during learning instead of adapting the fitting process. In our studied application of face landmark detection, we efficiently transfer facial expressions from other human individuals and are thus able to learn a precise face Active Appearance Model only from neutral faces of a single individual. Our approach is evaluated on two common face datasets and outperforms previous transfer methods.
Daniel Haase, Erik Rodner, Joachim Denzler
CVPR3
2014 Explorative Analysis of Heterogeneous, Unstructured, and Uncertain Data - A Computer Science Perspective on Biodiversity Research
abstract
We outline a blueprint for the development of new computer science approaches for the management and analysis of big data problems for biodiversity science. Such problems are characterized by a combination of different data sources each of which owns at least one of the typical characteristics of big data (volume, variety, velocity, or veracity). For these problems, we envision a solution that covers different aspects of integrating data sources and algorithms for their analysis on one of the following three layers: At the data layer, there are various data archives of heterogeneous, unstructured, and uncertain data. At the functional layer, the data are analyzed for each archive individually. At the meta-layer, multiple functional archives are combined for complex analysis.
Clemens Beckstein, Sebastian Böcker, Martin Bogdan, Helge Bruelheide, H. Martin Bücker, Joachim Denzler, Peter Dittrich, Ivo Grosse, Alexander Hinneburg, Birgitta König-Ries, Felicitas Löffler, Manja Marz, Matthias Müller-Hannemann, Wolf Zimmermann
DATA6
2014 Selecting Influential Examples: Active Learning with Expected Model Output Changes
Alexander Freytag, Erik Rodner, Joachim Denzler
ECCV (4)3
2014 A combination of generative and discriminative models for fast unsupervised activity recognition from traffic scene videos
abstract
Recent approaches in traffic and crowd scene analysis make extensive use of non-parametric hierarchical Bayesian models for intelligent clustering of features into activities. Although this has yielded impressive results, it requires the use of time consuming Bayesian inference during both training and classification. Therefore, we seek to limit Bayesian inference to the training stage, where unsupervised clustering is performed to extract semantically meaningful activities from the scene. In the testing stage, we use discriminative classifiers, taking advantage of their relative simplicity and fast inference. Experiments on publicly available data-sets show that our approach is comparable in classification accuracy to state-of-the-art methods and provides a significant speed-up in the testing phase.
Mahesh Venkata Krishna, Joachim Denzler
WACV2
2014 Intrinsic and extrinsic active self-calibration of multi-camera systems
Marcel Brückner, Ferid Bajramovic, Joachim Denzler
Mach. Vis. Appl.3
2013 Accurate 3D Multi-marker Tracking in X-ray Cardiac Sequences Using a Two-Stage Graph Modeling Approach
Daniel Haase, Marco Körner 0001, Wolfgang Bothe, Joachim Denzler
CAIP (2)5
2013 Temporal Self-Similarity for Appearance-Based Action Recognition in Multi-View Setups
Marco Körner 0001, Joachim Denzler
CAIP (1)2
2013 Kernel Null Space Methods for Novelty Detection
abstract
Detecting samples from previously unknown classes is a crucial task in object recognition, especially when dealing with real-world applications where the closed-world assumption does not hold. We present how to apply a null space method for novelty detection, which maps all training samples of one class to a single point. Beside the possibility of modeling a single class, we are able to treat multiple known classes jointly and to detect novelties for a set of classes with a single model. In contrast to modeling the support of each known class individually, our approach makes use of a projection in a joint subspace where training samples of all known classes have zero intra-class variance. This subspace is called the null space of the training data. To decide about novelty of a test sample, our null space approach allows for solely relying on a distance measure instead of performing density estimation directly. Therefore, we derive a simple yet powerful method for multi-class novelty detection, an important problem not studied sufficiently so far. Our novelty detection approach is assessed in comprehensive multi-class experiments using the publicly available datasets Caltech-256 and Image Net. The analysis reveals that our null space approach is perfectly suited for multi-class novelty detection since it outperforms all other methods.
Paul Bodesheim, Alexander Freytag, Erik Rodner, Michael Kemmler, Joachim Denzler
CVPR5
2013 Large-scale gaussian process multi-class classification for semantic segmentation and facade recognition
Björn Fröhlich, Erik Rodner, Michael Kemmler, Joachim Denzler
Mach. Vis. Appl.4
2013 One-class classification with Gaussian processes
Michael Kemmler, Erik Rodner, Esther-Sabrina Wacker, Joachim Denzler
Pattern Recognit.4
2013 Enhanced anomaly detection in wire ropes by combining structure and appearance
Esther-Sabrina Wacker, Joachim Denzler
Pattern Recognit. Lett.2
2012 Rapid Uncertainty Computation with Gaussian Processes and Histogram Intersection Kernels
Alexander Freytag, Erik Rodner, Paul Bodesheim, Joachim Denzler
ACCV (2)4
2012 Semantic Segmentation with Millions of Features: Integrating Multiple Cues in a Combined Random Forest Approach
Björn Fröhlich, Erik Rodner, Joachim Denzler
ACCV (1)3
2012 Analyzing the Subspaces Obtained by Dimensionality Reduction for Human Action Recognition from 3d Data
abstract
Since depth measuring devices for real-world scenarios became available in the recent past, the use of 3d data now comes more in focus of human action recognition. Due to the increased amount of data it seems to be advisable to model the trajectory of every landmark in the context of all other landmarks which is commonly done by dimensionality reduction techniques like PCA. In this paper we present an approach to directly use the subspaces (i.e. their basis vectors) for extracting features and classification of actions instead of projecting the landmark data themselves. This yields a fixed-length description of action sequences disregarding the number of provided frames. We give a comparison of various global techniques for dimensionality reduction and analyze their suitability for our proposed scheme. Experiments performed on the CMU Motion Capture dataset show promising recognition rates as well as robustness in the presence of noise and incorrect detection of landmarks.
Marco Körner 0001, Joachim Denzler
AVSS2
2012 Divergence-Based One-Class Classification Using Gaussian Processes
abstract
We present an information theoretic framework for one-class classification, which allows for deriving several new novelty scores. With these scores, we are able to rank samples according to their novelty and to detect outliers not belonging to a learnt data distribution. The key idea of our approach is to measure the impact of a test sample on the previously learnt model. This is carried out in a probabilistic manner using Jensen-Shannon divergence and reclassification results derived from the Gaussian pro-cess regression framework. Our method is evaluated using well-known machine learning datasets as well as large-scale image categorisation experiments showing its ability to achieve state-of-the-art performance. 1
Paul Bodesheim, Erik Rodner, Alexander Freytag, Joachim Denzler
BMVC4
2012 Large-Scale Gaussian Process Classification with Flexible Adaptive Histogram Kernels
Erik Rodner, Alexander Freytag, Paul Bodesheim, Joachim Denzler
ECCV (4)4
2012 Efficient semantic segmentation with Gaussian processes and histogram intersection kernels
Alexander Freytag, Björn Fröhlich, Erik Rodner, Joachim Denzler
ICPR4
2012 Finding discriminative features for Raman spectroscopy
Michael Kemmler, Joachim Denzler
ICPR2
2012 Scale-independent Spatio-temporal Statistical Shape Representations for 3D Human Action Recognition
Marco Körner 0001, Daniel Haase, Joachim Denzler
ICPRAM (1)3
2012 Tracking and Reconstruction in a Combined Optimization Approach
abstract
We present a novel approach to the structure-from-motion problem which combines the search for correspondences and geometric reconstruction, rather than treating these as separate steps. Through the combination of the two steps, we achieve an implicit feedback of 3D information to aid the correspondence search, and at the same time we avoid an explicit model for tracking errors. The reconstruction results are therefore optimal in case of, for example, Gaussian noise on image intensities. We also present an efficient online framework for structure-from-motion with our combined approach, thoroughly evaluate the method in experiments and compare the results to state-of-the-art methods.
Olaf Kähler, Joachim Denzler
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Combining Structure and Appearance for Anomaly Detection in Wire Ropes
Esther-Sabrina Wacker, Joachim Denzler
CAIP (2)2
2011 Exploiting the Manhattan-world assumption for extrinsic self-calibration of multi-modal sensor networks
abstract
Many new applications are enabled by combining a multi-camera system with a Time-of-Flight (ToF) camera, which is able to simultaneously record intensity and depth images. Classical approaches for self-calibration of a multi-camera system fail to calibrate such a system due to the very different image modalities. In addition, the typical environments of multi-camera systems are man-made and consist primary of only low textured objects. However, at the same time they satisfy the Manhattan-world assumption. We formulate the multi-modal sensor network calibration as a Maximum a Posteriori (MAP) problem and solve it by minimizing the corresponding energy function. First we estimate two separate 3D reconstructions of the environment: one using the pan-tilt unit mounted ToF camera and one using the multi-camera system. We exploit the Manhattan-world assumption and estimate multiple initial calibration hypotheses by registering the three dominant orientations of planes. These hypotheses are used as prior knowledge of a subsequent MAP estimation aiming to align edges that are parallel to these dominant directions. To our knowledge, this is the first self-calibration approach that is able to calibrate a ToF camera with a multi-camera system. Quantitative experiments on real data demonstrate the high accuracy of our approach.
Marcel Brückner, Joachim Denzler
ICCV2
2011 Learning with few examples for binary and multiclass classification using regularization of randomized trees
Erik Rodner, Joachim Denzler
Pattern Recognit. Lett.2
2010 One-Class Classification with Gaussian Processes
Michael Kemmler, Erik Rodner, Joachim Denzler
ACCV (2)3
2010 Multi-View Planning for Simultaneous Coverage and Accuracy Optimisation
abstract
S.118.1-118.11
Christoph Munkelt, Andreas Breitbarth, Gunther Notni, Joachim Denzler
BMVC4
2010 A Fast Approach for Pixelwise Labeling of Facade Images
abstract
Facade classification is an important subtask for automatically building large 3d city models. In the following we present an approach for pixel wise labeling of facade images using an efficient Randomized Decision Forest classifier and robust local color features. Experiments are performed with a popular facade dataset and a new demanding dataset of pixel wise labeled images from the Label Me project. Our method achieves high recognition rates and is significantly faster for training and testing than other Methods based on expensive feature transformation techniques.
Björn Fröhlich, Erik Rodner, Joachim Denzler
ICPR3
2010 Online Next-Best-View Planning for Accuracy Optimization Using an Extended E-Criterion
abstract
Next-best-view (NBV) planning is an important aspect for three-dimensional (3D) reconstruction within controlled environments, such as a camera mounted on a robotic arm. NBV methods aim at a purposive 3D reconstruction sustaining predefined goals and limitations. Up to now, literature mainly presents NBV methods for range sensors, model-based approaches or algorithms that address the reconstruction of a finite set of primitives. For this work, we use an intensity camera without active illumination. We present a novel combined online approach comprising feature tracking, 3D reconstruction, and NBV planning that addresses arbitrary unknown objects. In particular we focus on accuracy optimization based on the reconstruction uncertainty. To this end we introduce an extension of the statistical E-criterion to model directional uncertainty, and we present a closed-form, optimal solution to this NBV planning problem. Our experimental evaluation demonstrates the effectivity of our approach using an absolute error measure.
Michael Trummer, Christoph Munkelt, Joachim Denzler
ICPR3
2009 Combining Appearance and Range Based Information for Multi-class Generic Object Recognition
Doaa Hegazy, Joachim Denzler
CIARP2
2009 Randomized Probabilistic Latent Semantic Analysis for Scene Recognition
Erik Rodner, Joachim Denzler
CIARP2
2009 Geometric and probabilistic image dissimilarity measures for common field of view detection
abstract
Detecting image pairs with a common field of view is an important prerequisite for many computer vision tasks. Typically, common local features are used as a criterion for identifying such image pairs. This approach, however, requires a reliable method for matching features, which is generally a very difficult problem, especially in situations with a wide baseline or ambiguities in the scene. We propose two new approaches for the common field of view problem. The first one is still based on feature matching. Instead of requiring a very low false positive rate for the feature matching, however, geometric constraints are used to assess matches which may contain many false positives. The second approach completely avoids hard matching of features by evaluating the entropy of correspondence probabilities. We perform quantitative experiments on three different hand labeled scenes with varying difficulty. In moderately difficult situations with a medium baseline and few ambiguities in the scene, our proposed methods give similarly good results to the classical matching based method. On the most challenging scene having a wide baseline and many ambiguities, the performance of the classical method deteriorates, while ours are much less affected and still produce good results. Hence, our methods show the best overall performance in a combined evaluation.
Marcel Brückner, Ferid Bajramovic, Joachim Denzler
CVPR3
2009 Coarse registration of 3D surface triangulations based on moment invariants with applications to object alignment and identification
abstract
We present a new, direct way to register three-dimensional (3D) surfaces given the respective 3D points and surface triangulations. Our method is non-iterative and does not require any initial solution. The idea is to compute 3D invariants based on local surface moments. The resulting local surface descriptors are invariant with respect to Euclidean or to similarity transformations, by choice. In the final step we use the Hungarian method to find a minimum cost assignment of the computed descriptors. The method is robust against different point densities, noise and partial overlap. Our experiments with real data also show that the method can serve as automatic initialization of the iterative-closest-point (ICP) algorithm and, hence, extends the field of applications for this standard registration method.
Michael Trummer, Herbert Süße, Joachim Denzler
ICCV3
2009 Temporal Estimation of the 3d Guide-Wire Position Using 2d X-ray Images
Marcel Brückner, Frank Deinzer, Joachim Denzler
MICCAI (1)3
2009 A Framework for Actively Selecting Viewpoints in Object Recognition
abstract
Object recognition problems in computer vision are often based on single image data processing. In various applications this processing can be extended to a complete sequence of images, usually received passively. In contrast, we propose a method for active object recognition, where a camera is selectively moved around a considered object. Doing so, we aim at reliable classification results with a clearly reduced amount of necessary views by optimizing the camera movement for the access of new viewpoints (viewpoint selection). Therefore, the optimization criterion is the gain of class discriminative information when observing the appropriate next image. We show how to apply an unsupervised reinforcement learning algorithm to that problem. Specifically, we focus on the modeling of continuous states, continuous actions and supporting rewards for an optimized recognition. We also present an algorithm for the sequential fusion of gathered image information and we combine all these components into a single framework. The experimental evaluations are split into results for synthetic and real objects with one- or two-dimensional camera actions, respectively. This allows the systematic evaluation of the theoretical correctness as well as the practical applicability of the proposed method. Our experiments showed that the proposed combined viewpoint selection and viewpoint fusion approach is able to significantly improve the recognition rates compared to passive object recognition with randomly chosen views.
Frank Deinzer, Christian Derichs, Heinrich Niemann, Joachim Denzler
Int. J. Pattern Recognit. Artif. Intell.4
2008 Global Uncertainty-based Selection of Relative Poses for Multi Camera Calibration
abstract
Extrinsically calibrating a multi camera system from scene images is in gen-eral a very difficult problem. One promising approach uses pairwise relative poses as input. As a limited number of relative poses suffices, we propose automatically selecting only the most reliable ones. We present theoretically sound local and global uncertainty measures on relative poses and a selection criterion based on these measures. We show that our criterion is equivalent to computing a shortest subgraph consisting of shortest triangle connected paths and provide an efficient algorithm. In experiments on synthetic and real data, we show that our selection algorithm produces greatly improved calibration results, both, in case of varying portions of outliers as well as varying noise. 1
Ferid Bajramovic, Joachim Denzler
BMVC2
2008 Difference of Boxes Filters Revisited: Shadow Suppression and Efficient Character Segmentation
abstract
A robust segmentation is the most important part of anautomatic character recognition system (e.g. document processing, license plate recognition etc.). In our contribution we present an efficient segmentation framework using a preprocessing step for shadow suppression combined with a local thresholding technique. The method is based on a combination of difference of boxes filters and a new ternary segmentation, which are both simple low-level image operations. We also draw parallels to a recently published work on aganglion cell model and show that our approach is theoretically more substantiated as well as more robust and more efficient in practice. Systematic evaluation of noisy input data as well as results on a large dataset of license plate images show the robustness and efficiency of our proposed method. Our results can be applied easily to any optical character recognition system resulting in an impressive gain of robustness against nonlinear illumination.
Erik Rodner, Herbert Süße, Wolfgang Ortmann, Joachim Denzler
Document Analysis Systems4
2006 Aspects of Optimal Viewpoint Selection and Viewpoint Fusion
Frank Deinzer, Joachim Denzler, Christian Derichs, Heinrich Niemann
ACCV (2)2
2006 A Comparison of Nearest Neighbor Search Algorithms for Generic Object Recognition
Ferid Bajramovic, Frank Mattern, Nicholas Butko, Joachim Denzler
ACIVS4
2006 Integrated Viewpoint Fusion and Viewpoint Selection for Optimal Object Recognition
abstract
In the past decades, most object recognition systems were based on passive approaches. But in the last few years a lot of research was done in the field of active object recognition, that is selectively moving a sensor/camera around a considered object in order to acquire as much information about it as possible. In this paper we present an active object recognition approach that solves the problem of choosing optimal views (viewpoint selection) and iteratively fuses the gained information for an optimal 3D object recognition (viewpoint fusion) in an integrated manner. Therefore, we apply a method for the fusion of multiple views with respect to the knowledge about the assumed camera movement between them. For viewpoint selection we formally define the choice of additional views as an optimization problem. We show how to use reinforcement learning for this purpose and perform a training without user interaction. In this context we focus on the modeling of continuous states, continuous, one-dimensional actions and supporting rewards for an optimized recognition of real objects. The experimental results show that our combined viewpoint selection and viewpoint fusion approach is able to significantly improve the recognition rates compared to passive object recognition with randomly chosen views. 1
Frank Deinzer, Christian Derichs, Heinrich Niemann, Joachim Denzler
BMVC4
2005 Multi-step active object tracking with entropy based optimal actions using the sequential Kalman filter
abstract
We describe an enhanced method for the selection of optimal sensor actions in a probabilistic state estimation framework. We apply this to the selection of optimal focal lengths for cameras with a variable motor zoom in a real-time visual object tracking task. The optimal camera action is determined by the expected state estimate entropy for each candidate action. Varying action costs are taken into account by predicting the entropy several steps into the future. Our contribution is the use of the sequential Kalman filter to deal transparently with a variable number of cameras, potential object loss in a subset of the cameras, and to reduce the calculation time through independent optimization.
Benjamin Deutsch, Heinrich Niemann, Joachim Denzler
ICIP (3)3
2005 Appearance-based recognition of 3-D objects by cluttered background and occlusions
Michael Reinhold, Marcin Grzegorzek, Joachim Denzler, Heinrich Niemann
Pattern Recognit.3
2005 Markerless real-time 3-D target region tracking by motion backprojection from projection images
abstract
Accurate and fast localization of a predefined target region inside the patient is an important component of many image-guided therapy procedures. This problem is commonly solved by registration of intraoperative 2-D projection images to 3-D preoperative images. If the patient is not fixed during the intervention, the 2-D image acquisition is repeated several times during the procedure, and the registration problem can be cast instead as a 3-D tracking problem. To solve the 3-D problem, we propose in this paper to apply 2-D region tracking to first recover the components of the transformation that are in-plane to the projections. The 2-D motion estimates of all projections are backprojected into 3-D space, where they are then combined into a consistent estimate of the 3-D motion. We compare this method to intensity-based 2-D to 3-D registration and a combination of 2-D motion backprojection followed by a 2-D to 3-D registration stage. Using clinical data with a fiducial marker-based gold-standard transformation, we show that our method is capable of accurately tracking vertebral targets in 3-D from 2-D motion measured in X-ray projection images. Using a standard tracking algorithm (hyperplane tracking), tracking is achieved at video frame rates but fails relatively often (32% of all frames tracked with target registration error (TRE) better than 1.2 mm, 82% of all frames tracked with TRE better than 2.4 mm). With intensity-based 2-D to 2-D image registration using normalized mutual information (NMI) and pattern intensity (PI), accuracy and robustness are substantially improved. NMI tracked 82% of all frames in our data with TRE better than 1.2 mm and 96% of all frames with TRE better than 2.4 mm. This comes at the cost of a reduced frame rate, 1.7 s average processing time per frame and projection device. Results using PI were slightly more accurate, but required on average 5.4 s time per frame. These results are still substantially faster than 2-D to 3-D registration. We conclude that motion backprojection from 2-D motion tracking is an accurate and efficient method for tracking 3-D target motion, but tracking 2-D motion accurately and robustly remains a challenge.
Torsten Rohlfing, Joachim Denzler, Christoph Gräßl, Daniel B. Russakoff, Calvin R. Maurer Jr.
IEEE Trans. Medical Imaging2
2004 Active Sensing Strategies for Robotic Platforms, with an Application in Vision-Based Gripping
Benjamin Deutsch, Frank Deinzer, Matthias Zobel, Joachim Denzler
ICINCO (2)4
2004 Progressive Attenuation Fields: Fast 2D-3D Image Registration Without Precomputation
Torsten Rohlfing, Daniel B. Russakoff, Joachim Denzler, Calvin R. Maurer Jr.
MICCAI (1)3
2003 Viewpoint Selection - Planning Optimal Sequences of Views for Object Recognition
Frank Deinzer, Joachim Denzler, Heinrich Niemann
CAIP2
2003 Information Theoretic Focal Length Selection for Real-Time Active 3-D Object Tracking
abstract
Active object tracking, for example, in surveillance tasks, becomes more and more important these days. Besides the tracking algorithms themselves methodologies have to be developed for reasonable active control of the degrees of freedom of all involved cameras. We present an information theoretic approach that allows the optimal selection of the focal lengths of two cameras during active 3D object tracking. The selection is based on the uncertainty in the 3D estimation. This allows us to resolve the trade-off between small and large focal length: in the former case, the chance is increased to keep the object in the field of view of the cameras. In the latter one, 3D estimation becomes more reliable. Also, more details are provided, for example for recognizing the objects. Beyond a rigorous mathematical framework we present real-time experiments demonstrating that we gain an improvement in 3D trajectory estimation by up to 42% in comparison with tracking using a fixed focal length.
Joachim Denzler, Matthias Zobel, Heinrich Niemann
ICCV1
2003 MOBSY: Integration of vision and dialogue in service robots
Matthias Zobel, Joachim Denzler, Benno Heigl, Elmar Nöth, Dietrich Paulus, Jochen Schmidt, Georg Stemmer
Mach. Vis. Appl.2
2002 Entropy based camera control for visual object tracking
abstract
In active visual 3D object tracking, one goal is to control the pan and tilt axes of the involved cameras to keep the tracked object in the centers of the fields of view. We present a novel method, based on an information theoretic measure, that manages this task. The main advantage of the proposed approach is that there is no need for an explicit formulation of a camera controller, such as a PID-controller or something similar. For the case of Kalman filter based tracking, we demonstrate the practicability and evaluate the accuracy of the proposed method in simulations as well as in real-time tracking experiments.
Joachim Denzler, Matthias Zobel, Heinrich Niemann
ICIP (3)1
2002 Information Theoretic Sensor Data Selection for Active Object Recognition and State Estimation
abstract
We introduce a formalism for optimal sensor parameter selection for iterative state estimation in static systems. Our optimality criterion is the reduction of uncertainty in the state estimation process, rather than an estimator-specific metric (e.g., minimum mean squared estimate error). The claim is that state estimation becomes more reliable if the uncertainty and ambiguity in the estimation process can be reduced. We use Shannon's information theory to select information-gathering actions that maximize mutual information, thus optimizing the information that the data conveys about the true state of the system. The technique explicitly takes into account the a priori probabilities governing the computation of the mutual information. Thus, a sequential decision process can be formed by treating the a priori probability at a certain time step in the decision process as the a posteriori probability of the previous time step. We demonstrate the benefits of our approach in an object recognition application using an active camera for sequential gaze control and viewpoint selection. We describe experiments with discrete and continuous density representations that suggest the effectiveness of the approach.
Joachim Denzler
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 Robust Facial Feature Localization by Coupled Features
abstract
We consider the problem of robust localization of faces and some of their facial features. The task arises, e.g., in the medical field of visual analysis of facial paresis. We detect faces and facial features by means of appropriate DCT coefficients that we obtain by neatly using the coding capabilities of a JPEG hardware compressor. Beside an anthropometric localization approach we focus on how spatial coupling of the facial features can be used to improve robustness of the localization. Because the presented approach is embedded in a completely probabilistic framework, it is not restricted to facial features, it can be generalized to multipart objects of any kind. Therefore the notion of a "coupled structure" is introduced. Finally, the approach is applied to the problem of localizing facial features in DCT-coded images and results from our experiments are shown.
Matthias Zobel, Arnd Gebhard, Dietrich Paulus, Joachim Denzler, Heinrich Niemann
FG4
1999 Active Knowledge-Based Scene Analysis
Dietrich Paulus, Ulrike Ahlrichs, Benno Heigl, Joachim Denzler, Joachim Hornegger, Heinrich Niemann
ICVS4
1998 An Efficient Combination of 2D and 3D Shape Descriptions for Contour Based Tracking of Moving Objects
Joachim Denzler, Benno Heigl, Heinrich Niemann
ECCV (1)1
1997 Real-Time Pedestrian Tracking in Natural Scenes
Joachim Denzler, Heinrich Niemann
CAIP1
1997 Model Based Extraction of Articulated Objects in Image Sequences for Gait Analysis
abstract
This paper describes an approach to the extraction of articulated objects which will be used for gait analysis. In most medical applications markers are used to determine trajectories of different body parts. This approach works without any markers. Monotony operators which compute the displacement vector field are used to initialize a contour based tracking algorithm called active rays-for several body parts which are important for gait analysis. The contours of different parts of the human body are extracted and tracked. These parts are approached by simple 3D geometric objects (blocks), which 3D position and motion are estimated for the each image of the image sequence. Then, the trajectories of the moving parts represented by the 3D blocks can be determined and used for classification of different gait disorders.
Dorthe Meyer, Joachim Denzler, Heinrich Niemann
ICIP (3)2
1996 Statistical approach to classification of flow patterns for motion detection
abstract
We present a new approach for egomotion computation and the detection of independent motion in the scene. In contrast to related work we apply statistical methods which are based on the normal optical flow field. We extract features for supervised and unsupervised training from the normal optical flow field in order to train a Gaussian-distribution classifier (GDC) and a Kohonen feature map. Finally, in a test phase the egomotion computation is done by classifying features extracted from the normal optical flow field into the unknown motion direction. For the detection of independent motion, the scene is divided into regions. For each region a decision is made, whether the normal flow in this region is based on the camera motion or an independently moving object. We present results of this approach which show a recognition rate of up to 97% for the egomotion classification and a detection rate of moving objects of up to 87%.
Joachim Denzler, Volker Schless, Dietrich Paulus, Heinrich Niemann
ICIP (1)1
1996 3D data driven prediction for active contour models based on geometric bounding volumes
Joachim Denzler, Heinrich Niemann
Pattern Recognit. Lett.1
1994 Active Motion Detection and Object Tracking
abstract
In this paper we describe a two stage active vision system for tracking of a moving object which is detected in an overview image of the scene; a close-up view is then taken by changing the frame grabber's parameters and by a positional change of the camera mounted on a robot's hand. With a combination of several simple and fast working vision modules, a robust system for object tracking is constructed. The main principle is the use of two stages for object tracking: one for the detection of motion and one for the tracking itself. Errors in both stages can be detected in real time; then, the system switches back from the tracking to the motion detection stage. Standard UNIX interprocess communication mechanisms are used for the communication between control and vision modules. Object-oriented programming hides hardware details.>
Joachim Denzler, Dietrich Paulus
ICIP (3)1
1994 Learning, tracking and recognition of 3D objects
abstract
In this contribution we describe steps towards the implementation of an active robot vision system. In a sequence of images taken by a camera mounted on the hand of a robot, we detect, track, and estimate the position and orientation (pose) of a three-dimensional moving object. The extraction of the region of interest is done automatically by a motion tracking step. For learning 3-D objects using two-dimensional views and estimating the object's pose, a uniform statistical method is presented which is based on the expectation-maximization-algorithm (EM-algorithm). An explicit matching between features of several views is not necessary. The acquisition of the training sequence required for the statistical learning process needs the correlation between the image of an object and its pose; this is performed automatically by the robot. The robot's camera parameters are determined by a hand/eye-calibration and a subsequent computation of the camera position using the robot position. During the motion estimation stage the moving object is computed using active, elastic contours (snakes). We introduce a new approach for online initializing the snake on the first images of the given sequence, and show that the method of snakes is suited for real time motion tracking.>
Joachim Denzler, Rüdiger Bess, Joachim Hornegger, Heinrich Niemann, Dietrich Paulus
IROS1
1993 Going back to the source: inverse filtering of the speech signal with ANNs
abstract
In this paper we present a new method transforming speech signals to voice source signals (VSS) using articial neural networks (ANN).We will point out that the ANN mapping of speech signals into source signals is quite accurate, and most of the irregularities in the speech signal will lead to an irregularity in the source signal, produced by the ANN (ANN-VSS).We will show that the mapping of the ANN is robust with respect to untrained speakers, di erent recording conditions and facilities, and di erent v ocabularies.We will also present preliminary results which show that from the ANN source signal pitch periods can be determined accurately.
Joachim Denzler, Ralf Kompe, Andreas Kießling 0001, Heinrich Niemann, Elmar Nöth
EUROSPEECH1