Philipp Terhörst

dblp:203/2506 · DBLP profile ↗
← Back
36ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0001-8250-5712ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 23 · 8 first-author · 15 since 2021Security and privacy · 12 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author
YearPublicationVenuePosition
2026 Bay-CoFE: Bayesian consistency-driven feature elimination for eXplainable AI
abstract
Feature selection is a critical aspect of eXplainable Artificial Intelligence (XAI), and it has implications for model interpretability and predictive performance. CoFE (Consistency-driven Feature Elimination) framework was introduced recently using a frequentist approach. CoFE eliminates features with inconsistent coefficient signs in Linear Regression models by estimating the Sign Entropy (variability of the sign) of the coefficients using Bootstrapping. However, the uncertainty associated with estimating Sign Entropy using bootstrapping leads to slower convergence and inconsistency in feature subset selection in CoFE. In this paper, we present Bay-CoFE, 1 a Bayesian reformulation of CoFE, to solve the slow convergence and inconsistency issues of CoFE while retaining the benefits of selecting features with lower Sign Entropy in a Bayesian framework. We provide theoretical justifications and empirical evidence to prove Bay-CoFE’s superior convergence properties. Across all datasets, Bay-CoFE achieves significantly superior sign stability compared to traditional feature selection methods (Mann-Whitney U test, p-value < = 1.63e-03 and mean Cliff’s delta ≈ 0.83), with minimal predictive performance differences (Mann-Whitney U test, p-value > 0.1 and mean Cliff’s delta ≈ 0.32), demonstrating a highly favorable trade-off for interpretable modeling.
Revoti Prasad Bora, Philipp Terhörst, Raymond N. J. Veldhuis, Ramachandra Raghavendra, Kiran B. Raja
Neurocomputing2
2026 FRIES: Framework for inconsistency estimation of saliency metrics
abstract
Saliency maps are widely used as a post-hoc approach to explain the decision-making process of Deep Learning (DL) based image classification models, but evaluating their fidelity remains a complex problem. While saliency metrics have been introduced to evaluate the fidelity of saliency maps, existing saliency metrics, such as perturbation-based saliency metrics, have been previously reported to demonstrate statistical inconsistency. Although inconsistencies have been noted in different works, there exists no mechanism for estimating the same, i.e., Inconsistency Estimation (IE). Our primary objective is to address this limitation, and therefore, we propose a framework to estimate the inconsistency of saliency metrics for any given DL model. The framework enables building IE models for estimating the inconsistency by employing a set of perturbation types and schemes. The framework’s modular architecture provides flexibility across (i) perturbation types (Inpainting, Uniform, and Gaussian blur), (ii) perturbation schemes (pixel-wise and patch-wise), (iii) learning mechanisms (Convolutional Neural Networks and Vision Transformers) and (iv) IE modeling techniques (bagging and boosting). Extensive experimental results are shown on three well-known DL architectures (Inception-V3, Xception, and ResNet-50) on three different public datasets, including the Imagenette, Oxford-IIIT Pets Dataset, and PASCAL VOC 2007, along with results on ViTs for Oxford-IIIT Pets Dataset, and PASCAL VOC 2007. With a comprehensive evaluation of seven different perturbation types that include two inpainting, two Gaussian blur (with kernel widths of 0.9 and 1.5), and three uniform perturbations, our work shows the effectiveness of the proposed approach in estimating inconsistency. Statistically founded tests such as repeated cross-validation and the Permutation Test further validate the idea of the proposed framework for estimating the inconsistency of saliency metrics across unseen perturbations, making it useful in real-world scenarios.
Revoti Prasad Bora, Philipp Terhörst, Raymond N. J. Veldhuis, Ramachandra Raghavendra, Kiran B. Raja
Pattern Recognit.2
2025 Chasing Shadows: Solving Deepfake Detection Benchmarks Using Irrelevant Features Only
abstract
The emergence of Deepfake technology poses significant threats, particularly regarding misinformation and privacy. To mitigate these threats, Deepfake benchmarks play an important role in developing and testing reliable Deepfake detection algorithms. Consequently, it is crucial that these benchmarks do not possess serious biases that hinder the robustness and generalizability of Deepfake detectors during training and subsequently distort their true reliability during operation. This work investigates inherent biases in various Deepfake detection benchmark datasets by training simple classification models based on soft-biometric facial properties that do not contain Deepfake-related clues, i.e., decoy features. These mirage models reach up to 87.42% (balanced) accuracy on benchmark datasets using irrelevant decoy features alone for this task. As large parts of the performance of state-of-the-art models could also be achieved through exploiting benchmark biases, this raises the question of the unbiased performance of Deepfake detectors and their general reliability. Our analysis includes various Deepfake detection benchmarks and analyzes soft-biometric properties in determining their contribution to “solving” these benchmarks. Our findings underscore the need for more unbiased benchmarks beyond simply balancing demographic groups to enable future work on developing reliable solutions.
Philipp Terhörst, Marius Pedersen, Kiran B. Raja
FG2
2025 A Comprehensive Re-Evaluation of Biometric Modality Properties in the Modern Era
abstract
The rapid advancement of authentication systems and their increasing reliance on biometrics for faster and more accurate user verification experience, highlight the critical need for a reliable framework to evaluate the suitability of biometric modalities for specific applications. Currently, the most widely known evaluation framework is a comparative table from 1998, which no longer adequately captures recent technological developments or emerging vulnerabilities in biometric systems. To address these challenges, this work revisits the evaluation of biometric modalities through an expert survey involving 24 biometric specialists. The findings indicate substantial shifts in property ratings across modalities. For example, face recognition, shows improved ratings due to technological progress, while fingerprint, shows decreased reliability because of emerging vulnerabilities and attacks. Further analysis of expert agreement levels across rated properties highlighted the consistency of the provided evaluations and ensured the reliability of the ratings. Finally, expert assessments are compared with dataset-level uncertainty across 55 biometric datasets, revealing strong alignment in most modalities and underscoring the importance of integrating empirical evidence with expert insight. Moreover, the identified expert disagreements reveal key open challenges and help guide future research toward resolving them.
Rouqaiah Al-Refai, Pankaja Priya Ramasamy, Ragini Ramesh, Patricia Arias Cabarcos, Philipp Terhörst
IJCB5
2025 A Responsible Face Recognition Approach for Small and Mid-Scale Systems Through Personalized Neural Networks
abstract
Traditional face recognition systems rely on extracting fixed face representations, known as templates, to store and verify identities. These representations are typically generated by neural networks that often raise concerns regarding privacy, fairness, and explainability. In this work, we propose a novel model-template (MOTE) approach that replaces vector-based face templates with small personalized neural networks enabling more responsible face recognition for small and medium-scale systems. During enrollment, MOTE creates a dedicated binary classifier for each identity, trained to determine whether an input face matches the enrolled identity. Each classifier is trained using only a single reference sample, along with synthetically balanced samples to allow adjusting fairness at the level of a single individual during enrollment. Extensive experiments across multiple datasets and recognition systems demonstrate substantial improvements in privacy and fairness. Although the method increases inference time and storage requirements, it presents a strong solution for small- and mid-scale applications where privacy and fairness are critical. The code underlying our research is publicly available.1
Sebastian Gross, Stefan Heindorf, Philipp Terhörst
IJCB3
2025 BELIEF - Bayesian Sign Entropy Regularization for LIME Framework
abstract
Explanations of Local Interpretable Model-agnostic Explanations (LIME) are often inconsistent across different runs making them unreliable for eXplainable AI (XAI). The inconsistency stems from sign flips and variability in ranks of the segments for each different run. We propose a Bayesian Regularization approach to reduce sign flips, which in turn stabilizes feature rankings and ensures significantly higher consistency in explanations. The proposed approach enforces sparsity by incorporating a Sign Entropy prior on the coefficient distribution and dynamically eliminates features during optimization. Our results demonstrate that the explanations from the proposed method exhibit significantly better consistency and fidelity than LIME (and its earlier variants). Further, our approach exhibits comparable consistency and fidelity with a significantly lower execution time than the latest LIME variant, i.e., SLICE (CVPR 2024).
Revoti Prasad Bora, Philipp Terhörst, Raymond N. J. Veldhuis, Ramachandra Raghavendra, Kiran B. Raja
UAI2
2025 FALCON: Fair Face Recognition via Local Optimal Feature Normalization
abstract
Face recognition systems are widely used for identity verification in various fields. However, recent studies have highlighted bias issues related to demographic and non-demographic attributes such as accessories, haircolor, ethnicity, or gender. These biases lead to higher error rates for specific attribute subgroups. This is especially problematic in critical areas like forensics, where these systems are deployed. Addressing this issue requires a solution that reduces bias without compromising accuracy. Existing methods focus on learning less biased face representations, but they are often difficult to integrate into current systems or negatively impact overall recognition performance. This work introduces FALCON (Fair Adaptation Through Local Optimal Normalization), an effective method to increase fairness in face recognition systems. FALCON operates in an unsupervised manner, addressing bias without requiring demographic labels, and can be easily integrated as a post-processing step. It treats individuals with similar traits similarly, reducing bias in face recognition by processing each image individually. The proposed method is rigorously tested across various face recognition models and datasets, and compared with four other fairness post-processing methods. Results show that FALCON significantly enhances both fairness and accuracy. Unlike other methods, it allows seamlessly adjusting the fairness-accuracy trade-off while effectively addressing bias.
Rouqaiah Al-Refai, Philipp Hempel, Clara Biagi, Philipp Terhörst
WACV4
2025 Effective Backdoor Learning on Open-Set Face Recognition Systems
abstract
Backdoor attacks pose a serious threat to the security of face recognition systems. These involve the insertion of poisoned inputs into the training data to manipulate the model's behavior at inference time and can cause severe consequences, such as unauthorized access to secure systems or impersonation of legitimate users. Previous works on backdoor attacks have primarily focused on closed-set classification systems. However, open-set face recognition systems are commonly utilized in practical applications, which operate fundamentally differently from closed-set systems. In this paper, we propose two main contributions. First, we demonstrate that closed-set backdoor attacks are effective in basic classification scenarios but fail to perform well in complex open-set face recognition tasks. Second, we introduce Feature Stabilized Trigger Loss (FSTL), a novel loss function designed to facilitate the learning of backdoors in open-set recognition models. The experiments were conducted on two large-scale datasets using a variety of high-performing face recognition systems and by training with both physical and digital triggers. Since developing effective attack countermeasures requires knowledge of effective attacks, this work will enable future works on developing more secure recognition systems.
Diana Voth, Leonidas Dane, Jonas Grebe, Sebastian Peitz, Philipp Terhörst
WACV5
2024 SLICE: Stabilized LIME for Consistent Explanations for Image Classification
abstract
Local Interpretable Model-agnostic Explanations (LIME) - a widely used post-ad-hoc model agnostic ex-plainable AI (XAI) technique. It works by training a simple transparent (surrogate) model using random samples drawn around the neighborhood of the instance (image) to be explained (IE). Explanations are then extracted for a black-box model and a given IE, using the surrogate model. However, the explanations of LIME suffer from inconsistency across different runs for the same model and the same IE. We identify two main types of inconsistencies: variance in the sign and importance ranks of the segments (superpixels). These factors hinder LIME from obtaining consistent explanations. We analyze these inconsistencies and propose a new method, Stabilized LIME for Consistent Explanations (SLICE). The proposed method handles the stabilization problem in two aspects: using a novel feature selection technique to eliminate spurious superpixels and an adaptive perturbation technique to generate perturbed images in the neighborhood of IE. Our results demonstrate that the explanations from SLICE exhibit significantly better consistency and fidelity than LIME (and its variant BayLime).
Revoti Prasad Bora, Philipp Terhörst, Raymond N. J. Veldhuis, Ramachandra Raghavendra, Kiran B. Raja
CVPR2
2024 ASPECD: Adaptable Soft-Biometric Privacy-Enhancement Using Centroid Decoding for Face Verification
abstract
State-of-the-art face recognition models commonly extract information-rich biometric templates from the input images that are then used for comparison purposes and identity inference. While these templates encode identity information in a highly discriminative manner, they typically also capture other potentially sensitive facial attributes, such as age, gender or ethnicity. To address this issue, Soft-Biometric Privacy-Enhancing Techniques (SB-PETs) were proposed in the literature that aim to suppress such attribute information, and, in turn, alleviate the privacy risks associated with the extracted biometric templates. While various SB-PETs were presented so far, existing approaches do not provide dedicated mechanisms to determine which soft-biometrics to exclude and which to retain. In this paper, we address this gap and introduce ASPECD, a modular framework designed to selectively suppress binary and categorical soft-biometrics based on users' privacy preferences. ASPECD consists of multiple sequentially connected components, each dedicated for privacy-enhancement of an individual soft-biometric attribute. The proposed framework suppresses attribute information using a Moment-based Disentanglement process coupled with a centroid decoding procedure, ensuring that the privacy-enhanced templates are directly comparable to the templates in the original embedding space, regardless of the soft-biometric modality being suppressed. To validate the performance of ASPECD, we conduct experiments on a large-scale face dataset and with five state-of-the-art face recognition models, demonstrating the effectiveness of the proposed approach in suppressing single and multiple soft-biometric attributes. Our approach achieves a competitive privacy-utility trade-off compared to the state-of-the-art methods in scenarios that involve enhancing privacy w.r.t. gender and ethnicity attributes. The model will be made publicly available.
Peter Rot, Philipp Terhörst, Peter Peer, Vitomir Struc
FG2
2024 On the Trustworthiness of Face Morphing Attack Detectors
abstract
Morphing attacks blend face images of multiple distinct identities into a single photo combining their facial characteristics. Since the morphed image can be verified to multiple identities, these attacks greatly threaten face recognition systems. Morphing Attack Detection (MAD), addresses this problem by determining if an image is a bona fide or morphed image. Previous works on MAD focused on the models’ decision to improve their performance and generalizability. However, it has not been investigated whether the estimated probabilities of such decisions accurately reflect the true confidence with which predictions are made. Since these models struggle with unseen attacks, it is important to determine the models’ decision confidence to decide to trust or distrust the decisions. In this work, we (a) demonstrate that state-of-the-art MAD models struggle with providing reliable confidence estimations, (b) propose two new metrics to categorize their confidence prediction behavior, and (c) demonstrate that a simple calibration method can make the model more trustworthy without changing general model performance. The experiments were conducted in cross-dataset evaluation settings across two different MAD models using three different preprocessing and five morphing generation techniques. We found that both models are not calibrated and therefore not trustworthy. Moreover, it is shown that a simple and effective adjustment of the prediction confidence is possible. This will make future presentation attack detection, and especially MAD, models more trustworthy.
Rouqaiah Al-Refai, Clara Biagi, Kiran B. Raja, Ramachandra Raghavendra, Christoph Busch 0001, Philipp Terhörst
IJCB7
2024 CoFE: Consistency-Driven Feature Elimination for eXplainable AI
Revoti Prasad Bora, Philipp Terhörst, Raymond N. J. Veldhuis, Ramachandra Raghavendra, Kiran B. Raja
ICPR (9)2
2024 Efficient Explainable Face Verification based on Similarity Score Argument Backpropagation
abstract
Explainable Face Recognition is gaining growing attention as the use of the technology is gaining ground in security-critical applications. Understanding why two face images are matched or not matched by a given face recognition system is important to operators, users, and developers to increase trust, accountability, develop better systems, and highlight unfair behavior. In this work, we propose a similarity score argument backpropagation (xSSAB) approach that supports or opposes the face-matching decision to visualize spatial maps that indicate similar and dissimilar areas as interpreted by the underlying FR model. Furthermore, we present Patch-LFW, a new explainable face verification benchmark that enables along with a novel evaluation protocol, the first quantitative evaluation of the validity of similarity and dissimilarity maps in explainable face recognition approaches. We compare our efficient approach to state-of-the-art approaches demonstrating a superior trade-off between efficiency and performance. The code as well as the proposed Patch-LFW is publicly available at: https://github.com/marcohuber/xSSAB.
Marco Huber, Anh Thi Luu, Philipp Terhörst, Naser Damer
WACV3
2024 NeuroIDBench: An open-source benchmark framework for the standardization of methodology in brainwave-based authentication research
abstract
Biometric systems based on brain activity have been proposed as an alternative to passwords or to complement current authentication techniques. By leveraging the unique brainwave patterns of individuals, these systems offer the possibility of creating authentication solutions that are resistant to theft, hands-free, accessible, and potentially even revocable. However, despite the growing stream of research in this area, faster advance is hindered by reproducibility problems. Issues such as the lack of standard reporting schemes for performance results and system configuration, or the absence of common evaluation benchmarks, make comparability and proper assessment of different biometric solutions challenging. Further, barriers are erected to future work when, as so often, source code is not published open access. To bridge this gap, we introduce NeuroIDBench, a flexible open source tool to benchmark brainwave-based authentication models. It incorporates nine diverse datasets, implements a comprehensive set of pre-processing parameters and machine learning algorithms, enables testing under two common adversary models (known vs unknown attacker), and allows researchers to generate full performance reports and visualizations. We use NeuroIDBench to investigate the shallow classifiers and deep learning-based approaches proposed in the literature, and to test robustness across multiple sessions. We observe a 37.6% reduction in Equal Error Rate (EER) for unknown attacker scenarios (typically not tested in the literature), and we highlight the importance of session variability to brainwave authentication. All in all, our results demonstrate the viability and relevance of NeuroIDBench in streamlining fair comparisons of algorithms, thereby furthering the advancement of brainwave-based authentication through robust methodological practices.
Avinash Kumar Chaurasia, Matin Fallahi, Thorsten Strufe, Philipp Terhörst, Patricia Arias Cabarcos
J. Inf. Secur. Appl.4
2023 Explaining Face Recognition Through SHAP-Based Pixel-Level Face Image Quality Assessment
abstract
Biometric face recognition models are widely used in many different real-world applications. The output of these models can be used to make decisions that may strongly impact people. However, an explanation of how and why such outputs are derived is usually not given to humans. The lack of explainability of face recognition models leads to distrust in their decisions and does not encourage their use. The performance of face recognition models is influenced by the quality of the input image. In case the quality of a face image is too low, the face recognition system will reject it to avoid compromising its performance. The quality is evaluated by Face Image Quality (FIQ) approaches, which assigned quality scores to the input images. Pixel-level face image quality (PLFIQ) increases the explainability of quality scores by explaining face image quality at the pixel level. This allows the users of face recognition systems to spot low-quality areas and allows them to make guided corrections. Previous works introduced the concept of PLFIQ and proposed evaluation procedures. This work proposes a new way of computing PLFIQ values depending on given FIQ methods using Shapley Values. They score the contribution of each pixel to the overall image quality evaluation. Therefore, Integrating Shapley Values increases the explainability of the FIQ models. Results show that using these methods leads to significantly better and more robust PLFIQ values estimates and thus provide better explainability.
Clara Biagi, Louis Rethfeld, Arjan Kuijper, Philipp Terhörst
IJCB4
2023 QMagFace: Simple and Accurate Quality-Aware Face Recognition
abstract
In this work, we propose QMagFace, a simple and effective face recognition solution (QMagFace) that combines a quality-aware comparison score with a recognition model based on a magnitude-aware angular margin loss. The proposed approach includes model-specific face image qualities in the comparison process to enhance the recognition performance under unconstrained circumstances. Exploiting the linearity between the qualities and their comparison scores induced by the utilized loss, our quality-aware comparison function is simple and highly generalizable. The experiments conducted on several face recognition databases and benchmarks demonstrate that the introduced quality-awareness leads to consistent improvements in the recognition performance. Moreover, the proposed QMagFace approach performs especially well under challenging circumstances, such as cross-pose, cross-age, or cross-quality. Consequently, it leads to state-of-the-art performances on several face recognition benchmarks, such as 98.50% on AgeDB, 83.95% on XQLFQ, and 98.74% on CFP-FP. The code for QMagFace is publicly available1.
Philipp Terhörst, Malte Ihlefeld, Marco Huber, Naser Damer, Florian Kirchbuchner, Kiran B. Raja, Arjan Kuijper
WACV1
2022 Stating Comparison Score Uncertainty and Verification Decision Confidence Towards Transparent Face Recognition
Marco Huber, Philipp Terhörst, Florian Kirchbuchner, Naser Damer, Arjan Kuijper
BMVC2
2022 On the (Limited) Generalization of MasterFace Attacks and Its Relation to the Capacity of Face Representations
abstract
A MasterFace is a face image that can successfully match against a large portion of the population. Since their generation does not require access to the information of the enrolled subjects, MasterFace attacks represent a potential security risk for widely-used face recognition systems. Previous works proposed methods for generating such images and demonstrated that these attacks can strongly compromise face recognition. However, previous works followed evaluation settings consisting of older recognition models, limited cross-dataset and cross-model evaluations, and the use of low-scale testing data. This makes it hard to state the generalizability of these attacks. In this work, we comprehensively analyse the generalizability of MasterFace attacks in empirical and theoretical investigations. The empirical investigations include the use of six state-of-the-art face recognition models, cross-dataset and cross-model evaluation protocols, and utilizing testing datasets of significantly higher size and variance. The results indicate a low generalizability when MasterFaces are training on a different face recognition model than the one used for testing. In these cases, the attack performance is similar to zero-effort imposter attacks. In the theoretical investigations, we define and estimate the face capacity and the maximum MasterFace coverage under the assumption that identities in the face space are well separated. The current trend of increasing the fairness and generalizability in face recognition indicates that the vulnerability of future systems might further decrease. Future works might analyse the utility of MasterFaces for understanding and enhancing the robustness of face recognition models.
Philipp Terhörst, Florian Bierbaum, Marco Huber, Naser Damer, Florian Kirchbuchner, Kiran B. Raja, Arjan Kuijper
IJCB1
2022 Verification of Sitter Identity Across Historical Portrait Paintings by Confidence-aware Face Recognition
abstract
Verifying the identity of a person (sitter) portrayed in a historical painting is often a challenging but critical task in art historian research. In many cases, this information has been lost due to time or other circumstances and today there are only speculations of art historians about which person it could be. Art historians often use subjective factors for this purpose and then infer from the identity information about the person depicted in terms of his or her life, status, and era. On the other hand, automated face recognition has achieved a high level of accuracy, especially on photographs, and considers objective factors to determine the identity or verify a suspected identity. The limited amount of data, as well as the domain-specific challenges, make the use of automated face recognition methods in the domain of historic paintings difficult. We propose a specialized, likelihood-based fusion method to enable deep learning-based face recognition on historic portrait paintings. We additionally propose a method to accurately determine the confidence of the made decision to assist art historians in their research. For this purpose, we used a model trained on common photographs and adapted it to the domain of historical paintings through transfer learning. By using an underlying challenge dataset, we compute the likelihood for the assumed identity against reference images of the identity and fuse them to utilize as much information as possible. From these results of the likelihoods fusion, we then derive decision confidence to make statements to determine the certainty of the model’s decision. The experiments were carried out in a leave-one-out evaluation scenario on our created database, the largest authentic database of historic portrait paintings to date, consisting of over 760 portrait paintings of 210 different sitters by over 250 different artists. The experiments demonstrated, that a) the proposed approach outperforms pure face recognition solutions, b) the fusion approach effectively combines the sitter information towards a higher verification accuracy, and c) the proposed confidence estimation approach is highly successful in capturing the estimated accuracy of the decision. The meta-information of the used historic face images can be found at https://github.com/marcohuber/HistoricalFaces.
Marco Huber, Philipp Terhörst, Anh Thi Luu, Florian Kirchbuchner, Naser Damer
ICPR2
2021 MiDeCon: Unsupervised and Accurate Fingerprint and Minutia Quality Assessment based on Minutia Detection Confidence
abstract
An essential factor to achieve high accuracies in finger-print recognition systems is the quality of its samples. Previous works mainly proposed supervised solutions based on image properties that neglects the minutiae extraction process, despite that most fingerprint recognition techniques are based on detected minutiae. Consequently, a fingerprint image might be assigned a high quality even if the utilized minutia extractor produces unreliable information. In this work, we propose a novel concept of assessing minutia and fingerprint quality based on minutia detection confidence (MiDeCon). MiDeCon can be applied to an arbitrary deep learning based minutia extractor and does not require quality labels for learning. We propose using the detection reliability of the extracted minutia as its quality indicator. By combining the highest minutia qualities, MiDeCon also accurately determines the quality of a full fingerprint. Experiments are conducted on the publicly available databases of the FVC 2006 and compared against several baselines, such as NIST’s widely-used fingerprint image quality software NFIQ1 and NFIQ2. The results demonstrate a significantly stronger quality assessment performance of the proposed MiDeCon-qualities as related works on both, minutia- and fingerprint-level. The implementation is publicly available.
Philipp Terhörst, André Boller, Naser Damer, Florian Kirchbuchner, Arjan Kuijper
IJCB1
2021 Privacy-Enhancing Face Biometrics: A Comprehensive Survey
abstract
Biometric recognition technology has made significant advances over the last decade and is now used across a number of services and applications. However, this widespread deployment has also resulted in privacy concerns and evolving societal expectations about the appropriate use of the technology. For example, the ability to automatically extract age, gender, race, and health cues from biometric data has heightened concerns about privacy leakage. Face recognition technology, in particular, has been in the spotlight, and is now seen by many as posing a considerable risk to personal privacy. In response to these and similar concerns, researchers have intensified efforts towards developing techniques and computational models capable of ensuring privacy to individuals, while still facilitating the utility of face recognition technology in several application scenarios. These efforts have resulted in a multitude of privacy-enhancing techniques that aim at addressing privacy risks originating from biometric systems and providing technological solutions for legislative requirements set forth in privacy laws and regulations, such as GDPR. The goal of this overview paper is to provide a comprehensive introduction into privacy-related research in the area of biometrics and review existing work on Biometric Privacy-Enhancing Techniques (B-PETs) applied to face biometrics. To make this work useful for as wide of an audience as possible, several key topics are covered as well, including evaluation strategies used with B-PETs, existing datasets, relevant standards, and regulations and critical open issues that will have to be addressed in the future.
Blaz Meden, Peter Rot, Philipp Terhörst, Naser Damer, Arjan Kuijper, Walter J. Scheirer, Arun Ross, Peter Peer, Vitomir Struc
IEEE Trans. Inf. Forensics Secur.3
2021 MAAD-Face: A Massively Annotated Attribute Dataset for Face Images
abstract
Soft-biometrics play an important role in face biometrics and related fields since these might lead to biased performances, threaten the user's privacy, or are valuable for commercial aspects. Current face databases are specifically constructed for the development of face recognition applications. Consequently, these databases contain a large number of face images but lack in the number of attribute annotations and the overall annotation correctness. In this work, we propose a novel annotation-transfer pipeline that allows to accurately transfer attribute annotations from multiple source datasets to a target dataset. The transfer is based on a massive attribute classifier that can accurately state its prediction confidence. Using these prediction confidences, a high correctness of the transferred annotations is ensured. Applying this pipeline to the VGGFace2 database, we propose the MAAD-Face annotation database. It consists of 3.3M faces of over 9k individuals and provides 123.9M attribute annotations of 47 different binary attributes. Consequently, it provides 15 and 137 times more attribute annotations than CelebA and LFW. Our investigation on the annotation quality by three human evaluators demonstrated the superiority of the MAAD-Face annotations over existing databases. Additionally, we make use of the large number of high-quality annotations from MAAD-Face to study the viability of soft-biometrics for recognition, providing insights into which attributes support genuine and imposter decisions. The MAAD-Face annotations dataset is publicly available.
Philipp Terhörst, Daniel Fährmann, Jan Niklas Kolf, Naser Damer, Florian Kirchbuchner, Arjan Kuijper
IEEE Trans. Inf. Forensics Secur.1
2020 SER-FIQ: Unsupervised Estimation of Face Image Quality Based on Stochastic Embedding Robustness
abstract
Face image quality is an important factor to enable high-performance face recognition systems. Face quality assessment aims at estimating the suitability of a face image for the purpose of recognition. Previous work proposed supervised solutions that require artificially or human labelled quality values. However, both labelling mechanisms are error prone as they do not rely on a clear definition of quality and may not know the best characteristics for the utilized face recognition system. Avoiding the use of inaccurate quality labels, we proposed a novel concept to measure face quality based on an arbitrary face recognition model. By determining the embedding variations generated from random subnetworks of a face model, the robustness of a sample representation and thus, its quality is estimated. The experiments are conducted in a cross-database evaluation setting on three publicly available databases. We compare our proposed solution on two face embeddings against six state-of-the-art approaches from academia and industry. The results show that our unsupervised solution outperforms all other approaches in the majority of the investigated scenarios. In contrast to previous works, the proposed solution shows a stable performance over all scenarios. Utilizing the deployed face recognition model for our face quality assessment methodology avoids the training phase completely and further outperforms all baseline approaches by a large margin. Our solution can be easily integrated into current face recognition systems, and can be modified to other tasks beyond face recognition.
Philipp Terhörst, Jan Niklas Kolf, Naser Damer, Florian Kirchbuchner, Arjan Kuijper
CVPR1
2020 Learning privacy-enhancing face representations through feature disentanglement
abstract
Convolutional Neural Networks (CNNs) are today the de-facto standard for extracting compact and discriminative face representations (templates) from images in automatic face recognition systems. Due to the characteristics of CNN models, the generated representations typically encode a multitude of information ranging from identity to soft-biometric attributes, such as age, gender or ethnicity. However, since these representations were computed for the purpose of identity recognition only, the soft-biometric information contained in the templates represents a serious privacy risk. To mitigate this problem, we present in this paper a privacy-enhancing approach capable of suppressing potentially sensitive soft-biometric information in face representations without significantly compromising identity information. Specifically, we introduce a Privacy-Enhancing Face-Representation learning Network (PFRNet) that disentangles identity from attribute information in face representations and consequently allows to efficiently suppress soft-biometrics in face templates. We demonstrate the feasibility of PFRNet on the problem of gender suppression and show through rigorous experiments on the CelebA, Labeled Faces in the Wild (LFW) and Adience datasets that the proposed disentanglement-based approach is highly effective and improves significantly on the existing state-of-the-art.
Blaz Bortolato, Marija Ivanovska, Peter Rot, Janez Krizaj, Philipp Terhörst, Naser Damer, Peter Peer, Vitomir Struc
FG5
2020 Beyond Identity: What Information Is Stored in Biometric Face Templates?
abstract
Deeply-learned face representations enable the success of current face recognition systems. Despite the ability of these representations to encode the identity of an individual, recent works have shown that more information is stored within, such as demographics, image characteristics, and social traits. This threatens the user's privacy, since for many applications these templates are expected to be solely used for recognition purposes. Knowing the encoded information in face templates helps to develop bias-mitigating and privacy-preserving face recognition technologies. This work aims to support the development of these two branches by analysing face templates regarding 113 attributes. Experiments were conducted on two publicly available face embeddings. For evaluating the predictability of the attributes, we trained a massive attribute classifier that is additionally able to accurately state its prediction confidence. This allows us to make more sophisticated statements about the attribute predictability. The results demonstrate that up to 74 attributes can be accurately predicted from face templates. Especially non-permanent attributes, such as age, hairstyles, haircolors, beards, and various accessories, found to be easily-predictable. Since face recognition systems aim to be robust against these variations, future research might build on this work to develop more understandable privacy preserving solutions and build robust and fair face templates.
Philipp Terhörst, Daniel Fährmann, Naser Damer, Florian Kirchbuchner, Arjan Kuijper
IJCB1
2020 Face Quality Estimation and Its Correlation to Demographic and Non-Demographic Bias in Face Recognition
abstract
Face quality assessment aims at estimating the utility of a face image for the purpose of recognition. It is a key factor to achieve high face recognition performances. Currently, the high performance of these face recognition systems come with the cost of a strong bias against demographic and non-demographic sub-groups. Recent work has shown that face quality assessment algorithms should adapt to the deployed face recognition system, in order to achieve highly accurate and robust quality estimations. However, this could lead to a bias transfer towards the face quality assessment leading to discriminatory effects e.g. during enrolment. In this work, we present an in-depth analysis of the correlation between bias in face recognition and face quality assessment. Experiments were conducted on two publicly available datasets captured under controlled and uncontrolled circumstances with two popular face embeddings. We evaluated four state-of-the-art solutions for face quality assessment towards biases to pose, ethnicity, and age. The experiments showed that the face quality assessment solutions assign significantly lower quality values towards subgroups affected by the recognition bias demonstrating that these approaches are biased as well. This raises ethical questions towards fairness and discrimination which future works have to address.
Philipp Terhörst, Jan Niklas Kolf, Naser Damer, Florian Kirchbuchner, Arjan Kuijper
IJCB1
2020 Post-comparison mitigation of demographic bias in face recognition using fair score normalization
Philipp Terhörst, Jan Niklas Kolf, Naser Damer, Florian Kirchbuchner, Arjan Kuijper
Pattern Recognit. Lett.1
2019 Exploring the Channels of Multiple Color Spaces for Age and Gender Estimation from Face Images
Fadi Boutros, Naser Damer, Philipp Terhörst, Florian Kirchbuchner, Arjan Kuijper
FUSION3
2019 Multi-algorithmic Fusion for Reliable Age and Gender Estimation from Face Images
Philipp Terhörst, Marco Huber, Jan Niklas Kolf, Naser Damer, Florian Kirchbuchner, Arjan Kuijper
FUSION1
2019 Unsupervised privacy-enhancement of face representations using similarity-sensitive noise transformations
Philipp Terhörst, Naser Damer, Florian Kirchbuchner, Arjan Kuijper
Appl. Intell.1
2018 Minutiae-Based Gender Estimation for Full and Partial Fingerprints of Arbitrary Size and Shape
Philipp Terhörst, Naser Damer, Andreas Braun 0001, Arjan Kuijper
ACCV (1)1
2018 Fingerprint and Iris Multi-Biometric Data Indexing and Retrieval
abstract
Indexing of multi-biometric data is required to facilitate fast search in large-scale biometric systems. Previous works addressing this issue in multi-biometric databases focused on multi-instance indexing, mainly iris data. Few works addressed the indexing in multi-modal databases, with basic candidate list fusion solutions limited to joining face and fingerprint data. Iris and fingerprint are widely used in large-scale biometric systems where fast retrieval is a significant issue. This work proposes joint multi-biometric retrieval solution based on fingerprint and iris data. This solution is evaluated under eight different candidate list fusion approaches with variable complexity on a database of 10,000 reference and probe records of irises and fingerprints. Our proposed multi-biometric retrieval of fingerprint and iris data resulted in a reduction of the miss rate (1- hit rate) at 0.1% penetration rate by 93% compared to fingerprint indexing and 88% compared to iris indexing.
Naser Damer, Philipp Terhörst, Andreas Braun 0001, Arjan Kuijper
FUSION2
2018 Deep and Multi-Algorithmic Gender Classification of Single Fingerprint Minutiae
abstract
Accurate fingerprint gender estimation can positively affect several applications, since fingerprints are one of the most widely deployed biometrics. For example, gender classification in criminal investigations may significantly minimize the list of potential subjects. Previous work mainly offered solutions for the task of gender classification based on complete fingerprints. However, partial fingerprint captures are frequently occurring in many applications, including forensics and the fast growing field of consumer electronics. Moreover, partial fingerprints are not well-defined. Therefore, this work improves the gender decision performance on a well-defined partition of the fingerprint. It enhances gender estimation on the level of a single minutia. Working on this level, we propose three main contributions that were evaluated on a publicly available database. First, a convolutional neural network model is offered that outperformed baseline solutions based on hand crafted features. Second, several multi-algorithmic fusion approaches were tested by combining the outputs of different gender estimators that help further increase the classification accuracy. Third, we propose including minutia detection reliability in the fusion process, which leads to enhancing the total gender decision performance. The achieved gender classification performance of a single minutia is comparable to the accuracy that previous work reported on a quarter of aligned fingerprints including more than 25 minutiae.
Philipp Terhörst, Naser Damer, Andreas Braun 0001, Arjan Kuijper
FUSION1
2017 Indexing of Single and Multi-instance Iris Data Based on LSH-Forest and Rotation Invariant Representation
Naser Damer, Philipp Terhörst, Andreas Braun 0001, Arjan Kuijper
CAIP (2)2
2017 General borda count for multi-biometric retrieval
abstract
Indexing of multi-biometric data is required to facilitate fast search in large-scale biometric systems. Previous works addressing this issue were challenged by including biometric sources of different nature, utilizing the knowledge about the biometric sources, and optimizing and tuning the retrieval performance. This work presents a generalized multi-biometric retrieval approach that adapts the Borda count algorithm within an optimizable structure. The approach was tested on a database of 10k reference and probe instances of the left and the right irises. The experiments and comparisons to five baseline solutions proved to achieve advances in terms of general indexing performance, tunability to certain operating points, and response to missing data. A clear advantage of the proposed solution was noticed when faced by candidate lists of low quality.
Naser Damer, Philipp Terhörst, Andreas Braun 0001, Arjan Kuijper
IJCB2
2017 Efficient, Accurate, and Rotation-Invariant Iris Code
abstract
The large scale of the recently demanded biometric systems has put a pressure on creating a more efficient, accurate, and private biometric solutions. Iris biometrics is one of the most distinctive and widely used biometric characteristics. High-performing iris representations suffer from the curse of rotation inconsistency. This is usually solved by assuming a range of rotational errors and performing a number of comparisons over this range, which results in a high computational effort and limits indexing and template protection. This work presents a generic and parameter-free transformation of binary iris representation into a rotation-invariant space. The goal is to perform accurate and efficient comparison and enable further indexing and template protection deployment. The proposed approach was tested on a database of 10 000 subjects of the ISYN1 iris database generated by CASIA. Besides providing a compact and rotational-invariant representation, the proposed approach reduced the equal error rate by more than 55% and the computational time by a factor of up to 44 compared to the original representation.
Naser Damer, Philipp Terhörst, Andreas Braun 0001, Arjan Kuijper
IEEE Signal Process. Lett.2