Stefan Hörmann 0001

dblp:88/399-1 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
10since 2021 · last 2023
0000-0002-0086-340XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2023 Octuplet Loss: Make Face Recognition Robust to Image Resolution
abstract
Image resolution, or in general, image quality, plays an essential role in the performance of today's face recognition systems. To address this problem, we propose a novel combination of the popular triplet loss to improve robustness against image resolution via fine-tuning of existing face recognition models. With octuplet loss, we leverage the relationship between high-resolution images and their synthetically down-sampled variants jointly with their identity labels. Fine-tuning several state-of-the-art approaches with our method proves that we can significantly boost performance for cross-resolution (high-to-low resolution) face verification on various datasets without meaningfully exacerbating the performance on high-to-high resolution images. Our method applied on the FaceTransformer network achieves 95.12% face verification accuracy on the challenging XQLFW dataset while reaching 99.73% on the LFW database. Moreover, the low-to-low face verification accuracy benefits from our method. We release our code11Code available on https://github.com/Martlgap/octuplet-loss to allow seamless integration of the octuplet loss into existing frameworks.
Martin Knoche, Mohamed R. Elkadeem, Stefan Hörmann 0001, Gerhard Rigoll
FG3
2023 Frustration Recognition Using Spatio Temporal Data: A Novel Dataset and GCN Model to Recognize In-Vehicle Frustration
abstract
Frustration is an unpleasant emotion prevalent in several target applications of affective computing, such as human-machine interaction, learning, (online) customer interaction, and gaming. One idea to redeem this issue is to recognize frustration to offer help or mitigation in real-time, e.g., by a personal assistant. However, the recognition of frustration is not limited to these applied contexts but can also inform emotion research in general. This paper presents a dataset of 43 participants who experienced frustration in driving-related situations in a simulator. The data set contains a continuous subjective label, hand-annotated face and body expressions, facial landmark coordinates of two cameras, and the participants’ age and sex information. In addition, a descriptive analysis and description of the data's characteristics are provided together with a Graph Convolution Network based model to recognize frustration. Allowing for a tolerance of 10%, the model could correctly identify frustration with a similarity of 79.4 % and a variance of 7.7 %. This work is valuable for researchers of the affective computing community because it provides realistic data with an in-depth description of its characteristics and a benchmark model for automated frustration recognition. Our FRUST-dataset is publicly available under:https://ts.dlr.de/data-lake/frust-dataset/dataset.zip.
Esther Bosch, Raquel Le Houcq Corbí, Klas Ihme, Stefan Hörmann 0001, Meike Jipp, David Kaethner
IEEE Trans. Affect. Comput.4
2022 Dissected 3D CNNs: Temporal skip connections for efficient online video processing
Okan Köpüklü, Stefan Hörmann 0001, Fabian Herzog, Hakan Çevikalp, Gerhard Rigoll
Comput. Vis. Image Underst.2
2021 A Coarse-to-Fine Dual Attention Network for Blind Face Completion
abstract
In the area of face completion, the missing information within an occluded area is estimated, yielding a realistic face of the same identity. In most previous works, the mask describing the occluded region is known, limiting the scope of application. To alleviate this limitation, we propose a coarse-to-fine network trained as a conditional generative adversarial network. While the coarse network predicts the mask and generates a rough estimation of the semantic content, the subsequent fine network refines the rough prediction into a realistic and identity-persevering reconstruction. This is achieved by incorporating adversarial loss and using features from a pretrained face feature extractor. Unlike previous approaches, we employ two parallel attention mechanisms: 1) a patchwise cross-attention module to substitute information within the occluded patches with patches from the non-occluded region; 2) a pixel-wise global self-attention to allow information exchange within the entire feature map. Our exhaustive analysis, including reconstruction quality and face recognition metrics, shows that our approach outperforms the state of the art in blind face completion, improving the true positive identification rate at rank 1 on the MegaFace benchmark from 36.55 % to 42.48 %. This represents a substantial step towards closing the gap between occluded (29.34 %) and non-occluded faces (52.32 %). In terms of reconstruction quality, we obtain a structural similarity of 0.9639 compared to 0.8526 and 0.9563 for occluded faces and the state of the art, respectively. In addition to previous approaches, we provide an in-depth analysis of the influence of the position, size, and sparsity of the occlusion and use facial landmark prediction to measure reconstruction quality.
Stefan Hörmann 0001, Zhibing Xia, Martin Knoche, Gerhard Rigoll
FG1
2021 Cross-Quality LFW: A Database for Analyzing Cross- Resolution Image Face Recognition in Unconstrained Environments
abstract
Real-world face recognition applications often deal with suboptimal image quality or resolution due to different capturing conditions such as various subject-to-camera distances, poor camera settings, or motion blur. This characteristic has an unignorable effect on performance. Recent cross-resolution face recognition approaches used simple, arbitrary, and unrealistic down- and up-scaling techniques to measure robustness against real-world edge-cases in image quality. Thus, we propose a new standardized benchmark dataset and evaluation protocol derived from the famous Labeled Faces in the Wild (LFW). In contrast to previous derivatives, which focus on pose, age, similarity, and adversarial attacks, our Cross-Quality Labeled Faces in the Wild (XQLFW) maximizes the quality difference. It contains only more realistic synthetically degraded images when necessary. Our proposed dataset is then used to further investigate the influence of image quality on several state-of-the-art approaches. With XQLFW, we show that these models perform differently in cross-quality cases, and hence, the generalizing capability is not accurately predicted by their performance on LFW. Additionally, we report baseline accuracy with recent deep learning models explicitly trained for cross-resolution applications and evaluate the susceptibility to image quality. To encourage further research in cross-resolution face recognition and incite the assessment of image quality robustness, we publish the database and code for evaluation.11Code, dataset and evaluation protocol available on https://martlgap.github.io/xqlfw
Martin Knoche, Stefan Hörmann 0001, Gerhard Rigoll
FG2
2021 Lightweight Multi-Branch Network For Person Re-Identification
abstract
Person Re-Identification aims to retrieve person identities from images captured by multiple cameras or the same cameras in different time instances and locations. Because of its importance in many vision applications from surveillance to human-machine interaction, person re-identification methods need to be reliable and fast. While more and more deep architectures are proposed for increasing performance, those methods also increase overall model complexity. This paper proposes a lightweight network that combines global, part-based, and channel features in a unified multi-branch architecture that builds on the resource-efficient OSNet backbone. Using a well-founded combination of training techniques and design choices, our final model achieves state-of-the-art results on CUHK03 labeled, CUHK03 detected, and Market-1501 with 85.1% mAP/ 87.2% rankl, 82.4% mAP/84.9% rankl, and 91.5% mAP/96.3% rankl, respectively.
Fabian Herzog, Xunbo Ji, Torben Teepe, Stefan Hörmann 0001, Johannes Gilg, Gerhard Rigoll
ICIP4
2021 Face Texture Generation And Identity-Preserving Rectification
abstract
Textures are a vital asset in conveying a realistic impression of a 3D scene to the viewers. In order to obtain high-quality textures, real-life objects are scanned or designers create handcrafted textures. Both tasks involve manual work, are quite time-consuming, and therefore fail when a large quantity of textures is required. Thus, we propose to use a Generative Adversarial Network to generate an artificial texture. As textures need to be perfectly aligned with the 2D projection of the 3D model, our method involves a texture rectification technique, ensuring that the generated textures wrap well onto the 3D model. On the example of face textures, we illustrate that our method generates textures of high quality and variance. Moreover, we show that the rectification process preserves the facial appearance and identity, indicating that we successfully disentangle features responsible for facial appearance and the texture’s fit.
Stefan Hörmann 0001, Arka Bhowmick, Michael Weiher, Karl Leiss, Gerhard Rigoll
ICIP1
2021 Face Aggregation Network For Video Face Recognition
abstract
Typical approaches for video face recognition aggregate faces in a feature space to obtain a single feature representing the entire video. Unlike most previous approaches, we aggregate the faces directly in order to additionally obtain a single representative face as an intermediate output, from which a more discriminative feature vector is extracted. To overcome the limitation of a fixed number of input images of the state of the art in face aggregation, we incorporate a permutation invariant U-Net architecture capable of processing an arbitrary number of frames, which is employed in a generative adversarial network. We demonstrate the effectiveness of our method on three popular benchmark datasets for video face recognition. Our approach outperforms the baselines on the YouTube Faces dataset, obtaining an accuracy of 96.62%. Besides, we show that our method is robust against motion blur.
Stefan Hörmann 0001, Zhenxiang Cao, Martin Knoche, Fabian Herzog, Gerhard Rigoll
ICIP1
2021 Attention-Based Partial Face Recognition
abstract
Photos of faces captured in unconstrained environments, such as large crowds, still constitute challenges for current face recognition approaches as often faces are occluded by objects or people in the foreground. However, few studies have addressed the task of recognizing partial faces. In this paper, we propose a novel approach to partial face recognition capable of recognizing faces with different occluded areas. We achieve this by combining attentional pooling of a ResNet’s intermediate feature maps with a separate aggregation module. We further adapt common losses to partial faces in order to ensure that the attention maps are diverse and handle occluded parts. Our thorough analysis demonstrates that we outperform all baselines under multiple benchmark protocols, including naturally and synthetically occluded partial faces. This suggests that our method successfully focuses on the relevant parts of the occluded face.
Stefan Hörmann 0001, Martin Knoche, Torben Teepe, Gerhard Rigoll
ICIP1
2021 Gaitgraph: Graph Convolutional Network for Skeleton-Based Gait Recognition
abstract
Gait recognition is a promising video-based biometric for identifying individual walking patterns from a long distance. At present, most gait recognition methods use silhouette images to represent a person in each frame. However, silhouette images can lose fine-grained spatial information, and most papers do not regard how to obtain these silhouettes in complex scenes. Furthermore, silhouette images contain not only gait features but also other visual clues that can be recognized. Hence these approaches can not be considered as strict gait recognition. We leverage recent advances in human pose estimation to estimate robust skeleton poses directly from RGB images to bring back model-based gait recognition with a cleaner representation of gait. Thus, we propose GaitGraph that combines skeleton poses with Graph Convolutional Network (GCN) to obtain a modern model-based approach for gait recognition. The main advantages are a cleaner, more elegant extraction of the gait features and the ability to incorporate powerful spatiotemporal modeling using GCN. Experiments on the popular CASIA-B gait dataset show that our method archives state-of-the-art performance in model-based gait recognition.The code and models are publicly available1
Torben Teepe, Johannes Gilg, Fabian Herzog, Stefan Hörmann 0001, Gerhard Rigoll
ICIP5
2020 A Multi-Task Comparator Framework for Kinship Verification
abstract
Approaches for kinship verification often rely on cosine distances between face identification features. However, due to gender bias inherent in these features, it is hard to reliably predict whether two opposite-gender pairs are related. Instead of fine tuning the feature extractor network on kinship verification, we propose a comparator network to cope with this bias. After concatenating both features, cascaded local expert networks extract the information most relevant for their corresponding kinship relation. We demonstrate that our framework is robust against this gender bias and achieves comparable results on two tracks of the RFIW Challenge 2020. Moreover, we show how our framework can be further extended to handle partially known or unknown kinship relations.
Stefan Hörmann 0001, Martin Knoche, Gerhard Rigoll
FG1
2020 Attention Fusion for Audio-Visual Person Verification Using Multi-Scale Features
abstract
In the domain of audio-visual person recognition, many approaches use naive fusion techniques, such as scorelevel fusion or concatenation, to fuse the features obtained by face and audio extraction networks. More sophisticated methods fuse both features taking into account the quality of their corresponding inputs. In this paper, we propose a novel architecture to improve the prediction of feature quality. In contrary to previous works, which estimate feature quality based on the features themselves, we combine the information obtained from different layers of the feature extraction networks. In our analysis, we show that our approach outperforms state-of-the-art fusion approaches on well-established benchmarks for multimodal person verification. Moreover, we show that our model is robust against degradation of the visual input.
Stefan Hörmann 0001, Abdul Moiz, Martin Knoche, Gerhard Rigoll
FG1
2019 Gait Energy Image Restoration Using Generative Adversarial Networks
abstract
Gait is a biometric property that can be used for human identification in video surveillance. Basically, different gait features require motion of a person walking over one complete gait cycle. For example, in Gait Energy Image (GEI), average of silhouette images over one complete gait cycle is computed. However, in reality, there might be a partial gait cycle data available due to occlusion. In this paper, we propose a Generative Adversarial Network (GAN) in order to address the problem of gait recognition from incomplete gait cycle. Precisely, the network is able to reconstruct complete GEIs from incomplete GEIs. The proposed architecture is composed of (i) a generator which is an auto-encoder network to construct complete GEIs out of incomplete GEIs and (ii) two discriminators, one of which discriminates whether a given image is a full GEI while the other discriminates whether two GEIs belong to the same subject. We evaluate our approach on the OULP large gait dataset confirming that the proposed architecture successfully reconstructs complete GEIs from even extreme incomplete gait cycles.
Maryam Babaee, Okan Köpüklü, Stefan Hörmann 0001, Gerhard Rigoll
ICIP4
2019 Outlier-Robust Neural Aggregation Network for Video Face Identification
abstract
Current approaches for video face recognition rely on image sets containing faces of exclusively one identity. However, as image sets are created by unsupervised methods, it is necessary to consider outlier-afflicted sets for real-life applications. In this paper, we propose an Outlier-Robust Neural Aggregation Network (ORNAN). First, we embed each image into a feature space using a Convolutional Neural Network (CNN). With the help of two cascaded attention blocks, we predict outliers within the image set. By integrating this knowledge into our aggregation network, we adaptively aggregate all feature vectors to form a single feature, mitigating the influence of outliers and noisy features. We show that our network is robust against outliers using outlier-afflicted IJB-B and IJB-C benchmarks while maintaining similar performance without outliers.
Stefan Hörmann 0001, Martin Knoche, Maryam Babaee, Okan Köpüklü, Gerhard Rigoll
ICIP1
2019 Convolutional Neural Networks with Layer Reuse
abstract
A convolutional layer in a Convolutional Neural Network (CNN) consists of many filters which apply convolution operation to the input, capture some special patterns and pass the result to the next layer. If the same patterns also occur at the deeper layers of the network, why wouldn't the same convolutional filters be used also in those layers? In this paper, we propose a CNN architecture, Layer Reuse Network (LruNet), where the convolutional layers are used repeatedly without the need of introducing new layers to get a better performance. This approach introduces several advantages: (i) Considerable amount of parameters are saved since we are reusing the layers instead of introducing new layers, (ii) the Memory Access Cost (MAC) can be reduced since reused layer parameters can be fetched only once, (iii) the number of nonlinearities increases with layer reuse, and (iv) reused layers get gradient updates from multiple parts of the network. The proposed approach is evaluated on CIFAR-10, CIFAR-100 and Fashion-MNIST datasets for image classification task, and layer reuse improves the performance by 5.14%, 5.85% and 2.29%, respectively. The source code and pretrained models are publicly available1.
Okan Köpüklü, Maryam Babaee, Stefan Hörmann 0001, Gerhard Rigoll
ICIP3