EDBT 2026 Demo / reviewers in the wild / expert
Vitomir Struc
dblp:47/4796
· DBLP profile ↗
72ranked-venue papers
4as first author
48since 2021 · last 2026
0000-0002-3385-5780ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 2 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 4 first-author · 27 since 2021Security and privacy · 20 · 15 since 2021Human-computer interaction and ubiquitous computing · 17 · 13 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FunFace: Feature Utility and Norm Estimation for Face RecognitionabstractFace Recognition (FR) is used in a variety of application domains, from entertainment and; to security and surveillance. Such applications rely on the FR model to be robust and perform well in a variety of settings. To achieve this, state-of-the-art FR models typically use expressive adaptive margin loss functions, which tie the feature norm to concepts related to sample quality, such as recognizability and perceptual image quality. Recently, through the development of Face Image Quality Assessment (FIQA) techniques, biometric utility has become the preferred measure of face-image quality and has been shown to be a better predictor of the usefulness of samples for face recognition compared to more human-centric aspects, such as resolution, blur, and lighting, tied to general image quality. While image quality expressed through feature norms exhibits a certain level of correlation with biometric utility, it does not fully encapsulate all aspects of utility. To address this point, we propose a new adaptive margin loss, FunFace (Face Recognition Through Utility and Norm Estimation), which incorporates biometric utility, estimated by the Certainty Ratio, into the adaptive margin, taking inspiration from AdaFace. We show that FunFace (when used to train a face recognition model) achieves competitive results to other state-of-the-art FR models on benchmarks containing high-quality samples, while surpassing them on low quality benchmarks. The code is available at https://github.com/LSIbabnikz/FunFace. Ziga Babnik, Fadi Boutros, Naser Damer, Deepak Kumar Jain 0001, Peter Peer, Vitomir Struc |
FG | 6 |
| 2026 | IDSync: Improving diffusion models through identity classificationabstractNasl. z nasl. zaslona. Jernej Sabadin, Darian Tomasevic, Blaz Meden, Peter Peer, Vitomir Struc |
FG | 5 |
| 2026 | Employing Vision-Language Models for Face Image Quality Assessment
Erdi Saritas, Eren Onaran, Vitomir Struc, Hazim Kemal Ekenel |
FG | 3 |
| 2026 | EmoVisioNet: A hybrid network unifying lightweight CNN and attention-based vision model for facial emotion detectionabstractFacial emotion detection has witnessed a surge in demand across numerous applications, including human-computer interaction, healthcare, and security. Accurate expression recognition is crucial for improving human-computer interactions and understanding human behavior. Existing facial emotion detection models face challenges in achieving both high accuracy and real-time processing due to complex architectures. Our goal is to create an efficient yet accurate solution that can work on resource-constrained devices. To address the challenge of accurately recognizing emotions from facial expressions, we propose a novel hybrid approach that combines the strengths of pretrained Lightweight Convolutional Neural Networks (CNN), and Attention-based Vision Models. The pretrained Lightweight CNN serves as a feature extractor, efficiently capturing facial features, while the attention model refines the feature representation to focus on crucial regions of the face associated with different expressions. This enables our model to achieve state-of-the-art (SOTA) accuracy with reduced computational requirements. The proposed model, EmoVisioNet, achieves superior performance across multiple datasets, attaining 99.97 % accuracy on CK+, 96.23 % on RAF-DB, 93.88 % on FER2013, and 96.91 % on FERPlus. The obtained results surpass the current state-of-the-art in this field, demonstrating the EmoVisioNet’s superior performance in facial expression recognition. Gargi Mishra, Supriya Bajpai, Dharmender Saini, Rachna Jain, Deepak Kumar Jain 0001, Vitomir Struc |
Neurocomputing | 6 |
| 2026 | FaceMINT: A library for gaining insights into biometric face recognition via mechanistic interpretabilityabstractDeep-learning models, including those used in biometric recognition, have achieved remarkable performance on benchmark datasets as well as real-world recognition tasks. However, a major drawback of these models is their lack of transparency in decision-making. Mechanistic interpretability has emerged as a promising research field intended to help us gain insights into such models, but its application to biometric data remains limited. In this work, we bridge this gap by introducing the FaceMINT library, a publicly available Python library (build on top of Pytorch) that enables biometric researchers to inspect their models through mechanistic interpretability. It provides a plug-and-play solution that allows researchers to seamlessly switch between the analyzed biometric models, evaluate state-of-the-art sparse autoencoders, select from various image parametrizations, and fine-tune hyperparameters. Using a large scale Glint360K dataset, we demonstrate the usability of FaceMINT by applying its functionality to two state-of-the-art (deep-learning) face recognition models: AdaFace, based on Convolutional Neural Networks (CNN), and SwinFace, based on transformers. The proposed library implements various sparse auto-encoders (SAEs), including vanilla SAE, Gated SAE, JumpReLU SAE, and TopK SAE, which have achieved state-of-the-art results in the mechanistic interpretability of large language models. Our study highlights the promise of mechanistic interpretability in the biometric field, providing new avenues for researchers to explore model transparency and refine biometric recognition systems. The library is publicly available at www.gitlab.com/peterrot/facemint . Peter Rot, Robert Jutresa, Peter Peer, Vitomir Struc, Walter J. Scheirer, Klemen Grm |
Image Vis. Comput. | 4 |
| 2026 | Improving Emotion Recognition From Ambiguous Speech via Spatio-Temporal Spectrum Analysis and Real-Time Soft-Label CorrectionabstractSpeech represents a fundamental medium for conveying human emotions and, as a result, speech-based emotion recognition (SER) systems have become pivotal in advancing human-computer interaction (HCI) across a range of applications. While significant progress has been made in speech emotion recognition over recent years, existing solutions still face several key challenges, in that they:$(i)$rely excessively on subjectively annotated (discrete) labels during training,$(ii)$often overlook the label ambiguity of speech samples that express more than one class of emotions, and$(iii)$underutilize unlabeled or ambiguous speech, for which typically a label distribution (or so-called soft labels) is available. To address these issues, we propose in this paper a novel SER model that explicitly handles ambiguous speech samples and overcomes the shortcomings outlined above. Central to our approach is a novel real-time soft-label correction strategy designed to refine the annotations assigned to ambiguous speech. The proposed model leverages both, (explicitly) labeled as well as ambiguous samples and applies the dynamic soft-label correction strategy alongside an enhanced inter-class difference loss function to iteratively optimize the label distributions during training. We theoretically demonstrate that our method is capable of approximating the true emotional distribution of speech even in the presence of label noise, suggesting that utilizing ambiguous speech samples without explicit emotion labels still contributes toward more effective emotion recognition. Furthermore, we integrate the representational power of convolutional neural networks (CNNs) with the contextual modeling capabilities of Wav2Vec 2.0 to enable a comprehensive extraction of spatio-temporal speech features. Experimental results on the IEMOCAP multi-label dataset confirm the effectiveness of our approach, achieving state-of-the-art performance with significant improvements in weighted accuracy (WA) and unweighted accuracy (UA) over competing methods. Chenquan Gan, Daitao Zhou, Qingyi Zhu, Xibin Wang, Deepak Kumar Jain 0001, Vitomir Struc |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | SelfMAD: Enhancing Generalization and Robustness in Morphing Attack Detection via Self-Supervised LearningabstractWith the continuous advancement of generative models, face morphing attacks have become a significant challenge for existing face verification systems due to their potential use in identity fraud and other malicious activities. Contemporary Morphing Attack Detection (MAD) approaches frequently rely on supervised, discriminative models trained on examples of bona fide and morphed images. These models typically perform well with morphs generated with techniques seen during training, but often lead to sub-optimal performance when subjected to novel unseen morphing techniques. While unsupervised models have been shown to perform better in terms of generalizability, they typically result in higher error rates, as they struggle to effectively capture features of subtle artifacts. To address these shortcomings, we present SelfMAD, a novel self-supervised approach that simulates general morphing attack artifacts, allowing classifiers to learn generic and robust decision boundaries without overfitting to the specific artifacts induced by particular face morphing methods. Through extensive experiments on widely used datasets, we demonstrate that SelfMAD significantly outperforms current state-of-theart MADs, reducing the detection error by more than $\mathbf{6 4} \%$ in terms of EER when compared to the strongest unsupervised competitor, and by more than $66 \%$, when compared to the best performing discriminative MAD model, tested in crossmorph settings. The source code for SelfMAD is available at https://github.com/LeonTodorov/SelfMAD. Marija Ivanovska, Leon Todorov, Naser Damer, Deepak Kumar Jain 0001, Peter Peer, Vitomir Struc |
FG | 6 |
| 2025 | ID-Booth: Identity-consistent Face Generation with Diffusion ModelsabstractRecent advances in generative modeling have enabled the generation of high-quality synthetic data that is applicable in a variety of domains, including face recognition. Here, state-of-the-art generative models typically rely on conditioning and fine-tuning of powerful pretrained diffusion models to facilitate the synthesis of realistic images of a desired identity. Yet, these models often do not consider the identity of subjects during training, leading to poor consistency between generated and intended identities. In contrast, methods that employ identity-based training objectives tend to overfit on various aspects of the identity, and in turn, lower the diversity of images that can be generated. To address these issues, we present in this paper a novel generative diffusion-based framework, called ID-Booth. ID-Booth consists of a denoising network responsible for data generation, a variational auto-encoder for mapping images to and from a lower-dimensional latent space and a text encoder that allows for prompt-based control over the generation procedure. The framework utilizes a novel triplet identity training objective and enables identity-consistent image generation while retaining the synthesis capabilities of pretrained diffusion models. Experiments with a state-of-the-art latent diffusion model and diverse prompts reveal that our method facilitates better intra-identity consistency and inter-identity separability than competing methods, while achieving higher image diversity. In turn, the produced data allows for effective augmentation of small-scale datasets and training of betterperforming recognition models in a privacy-preserving manner. The source code for the ID-Booth framework is publicly available at https://github.com/dariant/ID-Booth. Darian Tomasevic, Fadi Boutros, Chenhao Lin, Naser Damer, Vitomir Struc, Peter Peer |
FG | 5 |
| 2025 | FROQ1: Observing Face Recognition Models for Efficient Quality AssessmentabstractFace Recognition (FR) plays a crucial role in many critical (high-stakes) applications, where errors in the recognition process can lead to serious consequences. Face Image Quality Assessment (FIQA) techniques enhance FR systems by providing quality estimates of face samples, enabling the systems to discard samples that are unsuitable for reliable recognition or lead to low-confidence recognition decisions. Most state-of-the-art FIQA techniques rely on extensive supervised training to achieve accurate quality estimation. In contrast, unsupervised techniques eliminate the need for additional training but tend to be slower and typically exhibit lower performance. In this paper, we introduce FROQ1(Face Recognition Observer of Quality), a semi-supervised, training-free approach that leverages specific intermediate representations within a given FR model to estimate face-image quality, and combines the efficiency of supervised FIQA models with the training-free approach of unsupervised methods. A simple calibration step based on pseudo-quality labels allows FROQ to uncover specific representations, useful for quality assessment, in any modern FR model. To generate these pseudo-labels, we propose a novel unsupervised FIQA technique based on sample perturbations. Comprehensive experiments with four state-of-the-art FR models and eight benchmark datasets show that FROQ leads to highly competitive results compared to the state-of-the-art, achieving both strong performance and efficient runtime, without requiring explicit training. The code for FROQ is available from: https://github.com/LSIbabnikz/FROQ Ziga Babnik, Deepak Kumar Jain 0001, Peter Peer, Vitomir Struc |
IJCB | 4 |
| 2025 | Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025abstractThis paper presents a summary of the 2025 Sclera Segmentation Benchmarking Competition (SSBC), which focused on the development of privacy-preserving sclera-segmentation models trained using synthetically generated ocular images. The goal of the competition was to evaluate how well models trained on synthetic data perform in comparison to those trained on real-world datasets. The competition featured two tracks: (i) one relying solely on synthetic data for model development, and (ii) one combining/mixing synthetic with (a limited amount of) real-world data. A total of nine research groups submitted diverse segmentation models, employing a variety of architectural designs, including transformer-based solutions, lightweight models, and segmentation networks guided by generative frameworks. Experiments were conducted across three evaluation datasets containing both synthetic and real-world images, collected under diverse conditions. Results show that models trained entirely on synthetic data can achieve competitive performance, particularly when dedicated training strategies are employed, as evidenced by the top performing models that achieved F1scores of over 0.8 in the synthetic data track. Moreover, performance gains in the mixed track were often driven more by methodological choices rather than by the inclusion of real data, highlighting the promise of synthetic data for privacy-aware biometric development. The code and data for the competition is available at: https://github.com/dariant/SSBC_2025. Matej Vitek, Darian Tomasevic, Abhijit Das 0001, Sabari Nathan, Gökhan Özbulak, G. A. T. Özbulak, Jean-Paul Calbimonte, André Anjos, Hariohm Hemant Bhatt, Dhruv Dhirendra Premani, Jay Chaudhari, Caiyong Wang, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Divya Velayudan, Maregu Assefa, Naoufel Werghi, Zachary A. Daniels, Leeon John, Ritesh Vyas, Jalil Nourmohammadi Khiarak, Taher Akbari Saeed, Mahsa Nasehi, Ali Kianfar, Mobina Pashazadeh Panahi, Geetanjali Sharma, Pushp Raj Panth, Ramachandra Raghavendra, Aditya Nigam, Umapada Pal 0001, Peter Peer, Vitomir Struc |
IJCB | 35 |
| 2025 | Optimizing ambiguous speech emotion recognition through spatial-temporal parallel network with label correction strategy
Chenquan Gan, Daitao Zhou, Qingyi Zhu, Deepak Kumar Jain 0001, Vitomir Struc |
Comput. Vis. Image Underst. | 6 |
| 2025 | Privacy-by-design AIoT vision for intelligent urban environmentsabstractThe recent advancements in AI (Artificial Intelligence) have been instrumental in fostering the development of AIoT (Artificial Intelligence of Things)-enabled urban environments. Machine vision and image analysis, in particular, have become integral to a wide array of AI applications within the field of urban planning and monitoring. Yet, the rapid adoption of AI algorithms in public areas has significantly heightened privacy concerns. In this paper, we introduce a privacy-by-design approach tailored for intelligent urban systems, presenting a holistic approach to the development and deployment of AI-driven systems for privacy-preserving image acquisition and analysis. Specifically, we design an embedded vision system that acquires privacy-protected data, safeguarding sensitive information against unauthorized access and potential misuse. Furthermore, we propose a strategy for developing AI vision methods using data that has been anonymized, ensuring that privacy is maintained throughout the AI application building process. Through experiments on a real-world AIoT-enabled urban environment use case - traffic flow monitoring at a city intersection - we demonstrate that our approach upholds strong privacy guarantees while maintaining the operational performance of modern AI vision systems. Marija Ivanovska, Jakob Kreft, Vitomir Struc, Janez Pers |
J. Syst. Archit. | 3 |
| 2025 | FICE: Text-conditioned fashion-image editing with guided GAN inversionabstractFashion-image editing is a challenging computer-vision task where the goal is to incorporate selected apparel into a given input image. Most existing techniques, known as Virtual Try-On methods, deal with this task by first selecting an example image of the desired apparel and then transferring the clothing onto the target person. Conversely, in this paper, we consider editing fashion images with text descriptions. Such an approach has several advantages over example-based virtual try-on techniques: (i) it does not require an image of the target fashion item, and (ii) it allows the expression of a wide variety of visual concepts through the use of natural language. Existing image-editing methods that work with language inputs are heavily constrained by their requirement for training sets with rich attribute annotations or they are only able to handle simple text descriptions. We address these constraints by proposing a novel text-conditioned editing model called FICE (Fashion Image CLIP Editing) that is capable of handling a wide variety of diverse text descriptions to guide the editing procedure. Specifically, with FICE, we extend the common GAN-inversion process by including semantic, pose-related, and image-level constraints when generating images. We leverage the capabilities of the CLIP model to enforce the text-provided semantics, due to its impressive image–text association capabilities. We furthermore propose a latent-code regularization technique that provides the means to better control the fidelity of the synthesized images. We validate the FICE through rigorous experiments on a combination of VITON images and Fashion-Gen text descriptions and in comparison with several state-of-the-art, text-conditioned, image-editing approaches. Experimental results demonstrate that the FICE generates very realistic fashion images and leads to better editing than existing, competing approaches. The source code is publicly available from: https://github.com/MartinPernus/FICE . • We propose a novel text-conditioned fashion image editing model called FICE. • FICE is an extended GAN inversion method tailored towards the fashion domain. • Descriptions are integrated with the GAN inversion process to manipulate appearance. • Comprehensive experiments show that FICE obtains state-of-the-art results. Martin Pernus, Clinton Fookes, Vitomir Struc, Simon Dobrisek |
Pattern Recognit. | 3 |
| 2025 | Hybrid Rumor Debunking in Online Social Networks: A Differential Game ApproachabstractOnline social networks (OSNs) facilitate the rapid and extensive spreading of rumors. While most existing methods for debunking rumors consider a solitary debunker, they overlook that rumor-mongering and debunking are interdependent and confrontational behaviors. In reality, a debunker must consider the impact of rumor-mongering behavior when making decisions. Moreover, a single rumor-debunking strategy is ineffective in addressing the complexity of the rumor environment in networks. Therefore, this article proposes a hybrid rumor-debunking approach that combines truth dissemination and regulatory measures based on the differential game theory under adversarial behaviors of rumor-mongering and debunking. Toward this end, we first establish a rumor propagation model using node-based modeling techniques that can be applied to any network structure. Next, we mathematically describe and analyze the processes of rumor-mongering and debunking. Finally, we validate the theoretical results of the proposed method through various comparative experiments, including comparisons with a random strategy, a uniform strategy, and single strategy models on real-world datasets collected from Facebook, Twitter, and YouTube. Furthermore, we harness two actual rumor events to estimate parameters and predict rumor propagation, thereby affirming the veracity and effectiveness of our rumor propagation model. Chenquan Gan, Wei Yang 0006, Qingyi Zhu, Deepak Kumar 0008, Vitomir Struc, Da-Wen Huang |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2024 | AdaDistill: Adaptive Knowledge Distillation for Deep Face Recognition
Fadi Boutros, Vitomir Struc, Naser Damer |
ECCV (55) | 2 |
| 2024 | WelcomeabstractIt was our pleasure and privilege to welcome you to Istanbul for the 18th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2024). We hope your experience at FG was rewarding both professionally and personally! Hazim Kemal Ekenel, Albert Ali Salah, Arun Ross, Vitomir Struc, Lale Akarun, Xilin Chen 0001, Shaun J. Canavan |
FG | 4 |
| 2024 | DiCTI: Diffusion-based Clothing Designer via Text-guided InputabstractRecent developments in deep generative models have opened up a wide range of opportunities for image synthesis, leading to significant changes in various creative fields, including the fashion industry. While numerous methods have been proposed to benefit buyers, particularly in virtual try-on applications, there has been relatively less focus on facilitating fast prototyping for designers and customers seeking to order new designs. To address this gap, we introduce DiCTI (Diffusion-based Clothing Designer via Text-guided Input), a straightforward yet highly effective approach that allows designers to quickly visualize fashion-related ideas using text inputs only. Given an image of a person and a description of the desired garments as input, DiCTI automatically generates multiple high-resolution, photorealistic images that capture the expressed semantics. By leveraging a powerful diffusion-based inpainting model conditioned on text inputs, DiCTI is able to synthesize convincing, high-quality images with varied clothing designs that viably follow the provided text descriptions, while being able to process very diverse and challenging inputs, captured in completely unconstrained settings. We evaluate DiCTI in comprehensive experiments on two different datasets (VITON-HD and Fashionpedia) and in comparison to the state-of-the-art (SoTa). The results of our experiments show that DiCTI convincingly outperforms the SoTA competitor in generating higher quality images with more elaborate garments and superior text prompt adherence, both according to standard quantitative evaluation measures and human ratings, generated as part of a user study. The source code of DiCTI will be made publicly available. Ajda Lampe, Julija Stopar, Deepak Kumar Jain 0001, Shinichiro Omachi, Peter Peer, Vitomir Struc |
FG | 6 |
| 2024 | ASPECD: Adaptable Soft-Biometric Privacy-Enhancement Using Centroid Decoding for Face VerificationabstractState-of-the-art face recognition models commonly extract information-rich biometric templates from the input images that are then used for comparison purposes and identity inference. While these templates encode identity information in a highly discriminative manner, they typically also capture other potentially sensitive facial attributes, such as age, gender or ethnicity. To address this issue, Soft-Biometric Privacy-Enhancing Techniques (SB-PETs) were proposed in the literature that aim to suppress such attribute information, and, in turn, alleviate the privacy risks associated with the extracted biometric templates. While various SB-PETs were presented so far, existing approaches do not provide dedicated mechanisms to determine which soft-biometrics to exclude and which to retain. In this paper, we address this gap and introduce ASPECD, a modular framework designed to selectively suppress binary and categorical soft-biometrics based on users' privacy preferences. ASPECD consists of multiple sequentially connected components, each dedicated for privacy-enhancement of an individual soft-biometric attribute. The proposed framework suppresses attribute information using a Moment-based Disentanglement process coupled with a centroid decoding procedure, ensuring that the privacy-enhanced templates are directly comparable to the templates in the original embedding space, regardless of the soft-biometric modality being suppressed. To validate the performance of ASPECD, we conduct experiments on a large-scale face dataset and with five state-of-the-art face recognition models, demonstrating the effectiveness of the proposed approach in suppressing single and multiple soft-biometric attributes. Our approach achieves a competitive privacy-utility trade-off compared to the state-of-the-art methods in scenarios that involve enhancing privacy w.r.t. gender and ethnicity attributes. The model will be made publicly available. Peter Rot, Philipp Terhörst, Peter Peer, Vitomir Struc |
FG | 4 |
| 2024 | Enhancing Gender Privacy with Photo-Realistic Fusion of Disentangled Spatial SegmentsabstractSoft-biometric privacy enhancing techniques (SB-PETs) transform facial images to preserve identity while preventing the automatic extraction of soft-biometrics by confusing machines through noise injections or attribute obfuscation. However, existing SB-PETs often sacrifice image quality for privacy enhancement, limiting practical usage, especially in applications that allow for human inspection. To address these issues, we introduce a novel SB-PET that (i) generates photo-realistic images with obscured gender information, which makes attribute extraction challenging for machine-learning models, but also human observers, and (ii) preserves identity to a significant extent. The proposed approach, abbreviated PriDSS, operates in the latent space of the StyleGANv2 model and aims to (i) preserve the appearance of facial parts from the input image carrying identity information, and (ii) incorporate global context from images of the opposite gender, thus, obscuring the original gender information. PriDSS shows promising results when compared to state-of-the-art techniques from the literature, and leads to competitive gender-privacy and face-verification performance, while ensuring superior photo-realism. Peter Rot, Janez Krizaj, Peter Peer, Vitomir Struc |
ICASSP | 4 |
| 2024 | Discovering Interpretable Feature Directions in the Embedding Space of Face Recognition ModelsabstractModern face recognition (FR) models, particularly their convolutional neural network based implementations, often raise concerns regarding privacy and ethics due to their "black-box" nature. To enhance the explainability of FR models and the interpretability of their embedding space, we introduce in this paper three novel techniques for discovering semantically meaningful feature directions (or axes). The first technique uses a dedicated facial-region blending procedure together with principal component analysis to discover embedding space direction that correspond to spatially isolated semantic face areas, providing a new perspective on facial feature interpretation. The other two proposed techniques exploit attribute labels to discern feature directions that correspond to intra-identity variations, such as pose, illumination angle, and expression, but do so either through a cluster analysis or a dedicated regression procedure. To validate the capabilities of the developed techniques, we utilize a powerful template decoder that inverts the image embedding back into the pixel space. Using the decoder, we visualize linear movements along the discovered directions, enabling a clearer understanding of the internal representations within face recognition models. The source code will be made publicly available. Richard Plesh, Janez Krizaj, Keivan Bahmani, Mahesh K. Banavar, Vitomir Struc, Stephanie Schuckers |
IJCB | 5 |
| 2024 | Deep Face Decoder: Towards understanding the embedding space of convolutional networks through visual reconstruction of deep face templatesabstractAdvances in deep learning and convolutional neural networks (ConvNets) have driven remarkable face recognition (FR) progress recently. However, the black-box nature of modern ConvNet-based face recognition models makes it challenging to interpret their decision-making process, to understand the reasoning behind specific success and failure cases, or to predict their responses to unseen data characteristics. It is, therefore, critical to design mechanisms that explain the inner workings of contemporary FR models and offer insight into their behavior. To address this challenge, we present in this paper a novel template-inversion approach capable of reconstructing high-fidelity face images from the embeddings (templates, feature-space representations) produced by modern FR techniques. Our approach is based on a novel Deep Face Decoder (DFD) trained in a regression setting to visualize the information encoded in the embedding space with the goal of fostering explainability. We utilize the developed DFD model in comprehensive experiments on multiple unconstrained face datasets, namely Visual Geometry Group Face dataset 2 (VGGFace2), Labeled Faces in the Wild (LFW), and Celebrity Faces Attributes Dataset High Quality (CelebA-HQ). Our analysis focuses on the embedding spaces of two distinct face recognition models with backbones based on the Visual Geometry Group 16-layer model (VGG-16) and the 50-layer Residual Network (ResNet-50). The results reveal how information is encoded in the two considered models and how perturbations in image appearance due to rotations, translations, scaling, occlusion, or adversarial attacks, are propagated into the embedding space. Our study offers researchers a deeper comprehension of the underlying mechanisms of ConvNet-based FR models, ultimately promoting advancements in model design and explainability. Janez Krizaj, Richard Plesh, Mahesh K. Banavar, Stephanie Schuckers, Vitomir Struc |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Generating bimodal privacy-preserving data for face recognitionabstractThe performance of state-of-the-art face recognition systems depends crucially on the availability of large-scale training datasets. However, increasing privacy concerns nowadays accompany the collection and distribution of biometric data, which has already resulted in the retraction of valuable face recognition datasets. The use of synthetic data represents a potential solution, however, the generation of privacy-preserving facial images useful for training recognition models is still an open problem. Generative methods also remain bound to the visible spectrum, despite the benefits that multispectral data can provide. To address these issues, we present a novel identity-conditioned generative framework capable of producing large-scale recognition datasets of visible and near-infrared privacy-preserving face images. The framework relies on a novel identity-conditioned dual-branch style-based generative adversarial network to enable the synthesis of aligned high-quality samples of identities determined by features of a pretrained recognition model. In addition, the framework incorporates a novel filter to prevent samples of privacy-breaching identities from reaching the generated datasets and improve both identity separability and intra-identity diversity. Extensive experiments on six publicly available datasets reveal that our framework achieves competitive synthesis capabilities while preserving the privacy of real-world subjects. The synthesized datasets also facilitate training more powerful recognition models than datasets generated by competing methods or even small-scale real-world datasets. Employing both visible and near-infrared data for training also results in higher recognition accuracy on real-world visible spectrum benchmarks. Therefore, training with multispectral data could potentially improve existing recognition systems that utilize only the visible spectrum, without the need for additional sensors. Darian Tomasevic, Fadi Boutros, Naser Damer, Peter Peer, Vitomir Struc |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Y-GAN: Learning dual data representations for anomaly detection in imagesabstractWe propose a novel reconstruction-based model for anomaly detection in image data, called’Y-GAN’. The model consists of a Y-shaped auto-encoder and represents images in two separate latent spaces. The first captures meaningful image semantics, which are key for representing (normal) training data, whereas the second encodes low-level residual image characteristics. To ensure the dual representations encode mutually exclusive information, a disentanglement procedure is designed around a latent (proxy) classifier. Additionally, a novel representation-consistency mechanism is proposed to prevent information leakage between the latent spaces. The model is trained in a one-class learning setting using only normal training data. Due to the separation of semantically-relevant and residual information, Y-GAN is able to derive informative data representations that allow for efficacious anomaly detection across a diverse set of anomaly detection tasks. The model is evaluated in comprehensive experiments with several recent anomaly detection models using four popular image datasets, i.e., MNIST, FMNIST, CIFAR10, and PlantVillage. Experimental results show that Y-GAN outperforms all tested models by a considerable margin and yields state-of-the-art results. The source code for the model is made publicly available at https://github.com/MIvanovska/Y-GAN. Marija Ivanovska, Vitomir Struc |
Expert Syst. Appl. | 2 |
| 2024 | A graph neural network with context filtering and feature correction for conversational emotion recognitionabstractConversational emotion recognition represents an important machine-learning problem with a wide variety of deployment possibilities. The key challenge in this area is how to properly capture the key conversational aspects that facilitate reliable emotion recognition, including utterance semantics, temporal order, informative contextual cues, speaker interactions as well as other relevant factors. In this paper, we present a novel Graph Neural Network approach for conversational emotion recognition at the utterance level. Our method addresses the outlined challenges and represents conversations in the form of graph structures that naturally encode temporal order, speaker dependencies, and even long-distance context. To efficiently capture the semantic content of the conversations, we leverage the zero-shot feature-extraction capabilities of pre-trained large-scale language models and then integrate two key contributions into the graph neural network to ensure competitive recognition results. The first is a novel context filter that establishes meaningful utterance dependencies for the graph construction procedure and removes low-relevance and uninformative utterances from being used as a source of contextual information for the recognition task. The second contribution is a feature-correction procedure that adjusts the information content in the generated feature representations through a gating mechanism to improve their discriminative power and reduce emotion-prediction errors. We conduct extensive experiments on four commonly used conversational datasets, i.e., IEMOCAP, MELD, Dailydialog, and EmoryNLP, to demonstrate the capabilities of the developed graph neural network with context filtering and error-correction capabilities. The results of the experiments point to highly promising performance, especially when compared to state-of-the-art competitors from the literature. Chenquan Gan, Jiahao Zheng 0007, Qingyi Zhu, Deepak Kumar Jain 0001, Vitomir Struc |
Inf. Sci. | 5 |
| 2024 | Transfer-learning enabled micro-expression recognition using dense connections and mixed attentionabstractMicro-expression recognition (MER) is a challenging computer vision problem, where the limited amount of available training data and insufficient intensity of the facial expressions are among the main issues adversely affecting the performance of existing recognition models. To address these challenges, this paper explores a transfer–learning enabled MER model using a densely connected feature extraction module with mixed attention. Unlike previous works that utilize transfer learning to facilitate MER and extract local facial-expression information, our model relies on pretraining with three diverse macro-expression datasets and, as a result, can: ( i ) overcome the problem of insufficient sample size and limited training data availability, ( i i ) leverage (related) domain-specific information from multiple datasets with diverse characteristics, and ( i i i ) improve the model adaptability to complex scenes. Furthermore, to enhance the intensity of the micro-expressions and improve the discriminability of the extracted features, the Euler video magnification (EVM) method is adopted in the preprocessing stage and then used jointly with a densely connected feature extraction module and a mixed attention mechanism to derive expressive feature representations for the classification procedure. The proposed feature extraction mechanism not only guarantees the integrity of the extracted features but also efficiently captures local texture cues by aggregating the most salient information from the generated feature maps, which is key for the MER task. The experimental results on multiple datasets demonstrate the robustness and effectiveness of our model compared to the state-of-the-art. Chenquan Gan, Qingyi Zhu, Deepak Kumar Jain 0001, Vitomir Struc |
Knowl. Based Syst. | 5 |
| 2024 | Fairness in face presentation attack detection
Meiling Fang, Wufei Yang, Arjan Kuijper, Vitomir Struc, Naser Damer |
Pattern Recognit. | 4 |
| 2024 | PrivacyProber: Assessment and Detection of Soft-Biometric Privacy-Enhancing TechniquesabstractSoft–biometric privacy–enhancing techniques represent machine learning methods that aim to: (i) mitigate privacy concerns associated with face recognition technology by suppressing selected soft–biometric attributes in facial images (e.g., gender, age, ethnicity) and (ii) make unsolicited extraction of sensitive personal information infeasible. Because such techniques are increasingly used in real–world applications, it is imperative to understand to what extent the privacy enhancement can be inverted and how much attribute information can be recovered from privacy–enhanced images. While these aspects are critical, they have not been investigated in the literature so far. In this paper, we, therefore,study the robustnessof several state–of–the–art soft–biometric privacy–enhancing techniques to attribute recovery attempts. We propose PrivacyProber, a high–level framework for restoring soft–biometric information from privacy–enhanced facial images, and apply it for attribute recovery in comprehensive experiments on three public face datasets, i.e., LFW, MUCT and Adience. Our experiments show that the proposed framework is able to restore a considerable amount of suppressed information, regardless of the privacy–enhancing technique used (e.g., adversarial perturbations, conditional synthesis, etc.), but also that there are significant differences between the considered privacy models. These results point to the need for novel mechanisms that can improve the robustness of existing privacy–enhancing techniques and secure them against potential adversaries trying to restore suppressed information. Additionally, we demonstrate that PrivacyProber can also be used to detect privacy–enhancement in facial images (under black–box assumptions) with high accuracy. Specifically, we show that a detection procedure can be developed around the proposed framework that islearning freeand, therefore, generalizes well across different data characteristics and privacy–enhancing techniques. Peter Rot, Klemen Grm, Peter Peer, Vitomir Struc |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | GlassesGAN: Eyewear Personalization Using Synthetic Appearance Discovery and Targeted Subspace ModelingabstractWe present GlassesGAN, a novel image editing frame-work for custom design of glasses, that sets a new standard in terms of output-image quality, edit realism, and continuous multi-style edit capability. To facilitate the editing process with GlassesGAN, we propose a Targeted Subspace Modelling (TSM) procedure that, based on a novel mechanism for (synthetic) appearance discovery in the latent space of a pre-trained GAN generator, constructs an eyeglasses-specific (latent) subspace that the editing framework can utilize. Additionally, we also introduce an appearance-constrained subspace initialization (SI) technique that centers the latent representation of the given input image in the well-defined part of the constructed sub-space to improve the reliability of the learned edits. We test GlassesGAN on two (diverse) high-resolution datasets (CelebA-HQ and SiblingsDB-HQf) and compare it to three state-of-the-art baselines, i.e., InterfaceGAN, GANSpace, and MaskGAN. The reported results show that GlassesGAN convincingly outperforms all competing techniques, while offering functionality (e.g., fine-grained multi-style editing) not available with any of the competitors. The source code for GlassesGAN is made publicly available. Richard Plesh, Peter Peer, Vitomir Struc |
CVPR | 3 |
| 2023 | DifFIQA: Face Image Quality Assessment Using Denoising Diffusion Probabilistic ModelsabstractModern face recognition (FR) models excel in constrained scenarios, but often suffer from decreased performance when deployed in unconstrained (real-world) environments due to uncertainties surrounding the quality of the captured facial data. Face image quality assessment (FIQA) techniques aim to mitigate these performance degradations by providing FR models with sample-quality predictions that can be used to reject low-quality samples and reduce false match errors. However, despite steady improvements, ensuring reliable quality estimates across facial images with diverse characteristics remains challenging. In this paper, we present a powerful new FIQA approach, named DifFIQA, which relies on denoising diffusion probabilistic models (DDPM) and ensures highly competitive results. The main idea behind the approach is to utilize the forward and backward processes of DDPMs to perturb facial images and quantify the impact of these perturbations on the corresponding image embeddings for quality prediction. Because the diffusion-based perturbations are computationally expensive, we also distill the knowledge encoded in DifFIQA into a regression-based quality predictor, called DifFIQA(R), that balances performance and execution time. We evaluate both models in comprehensive experiments on 7 diverse datasets, with 4 target FR models and against 10 state-of-the-art FIQA techniques with highly encouraging results. The source code is available from: https://github.com/LSIbabnikz/DifFIQA. Ziga Babnik, Peter Peer, Vitomir Struc |
IJCB | 3 |
| 2023 | Sclera Segmentation and Joint Recognition Benchmarking Competition: SSRBC 2023abstractThis paper presents the summary of the Sclera Segmentation and Joint Recognition Benchmarking Competition (SSRBC 2023) held in conjunction with IEEE International Joint Conference on Biometrics (IJCB 2023). Different from the previous editions of the competition, SSRBC 2023 not only explored the performance of the latest and most advanced sclera segmentation models, but also studied the impact of segmentation quality on recognition performance. Five groups took part in SSRBC 2023 and submitted a total of six segmentation models and one recognition technique for scoring. The submitted solutions included a wide variety of conceptually diverse deep-learning models and were rigorously tested on three publicly available datasets, i.e., MASD, SBVPI and MOBIUS. Most of the segmentation models achieved encouraging segmentation and recognition performance. Most importantly, we observed that better segmentation results always translate into better verification performance. Abhijit Das 0001, Saurabh Atreya, Aritra Mukherjee, Matej Vitek, Caiyong Wang, Guangzhe Zhao, Fadi Boutros, Patrick Siebke, Jan Niklas Kolf, Naser Damer, Sun Ye, Lu Hexin, Fan Aobo, You Sheng, Sabari Nathan, R. Suganya 0001, Rampriya Rajendran Shanthi, Geetanjali Sharma, P. Priyanka, Aditya Nigam, Peter Peer, Umapada Pal 0001, Vitomir Struc |
IJCB | 24 |
| 2023 | The Unconstrained Ear Recognition Challenge 2023: Maximizing Performance and Minimizing BiasabstractThe paper provides a summary of the 2023 Unconstrained Ear Recognition Challenge (UERC), a benchmarking effort focused on ear recognition from images acquired in uncontrolled environments. The objective of the challenge was to evaluate the effectiveness of current ear recognition techniques on a challenging ear dataset while analyzing the techniques from two distinct aspects, i.e., verification performance and bias with respect to specific demographic factors, i.e., gender and ethnicity. Seven research groups participated in the challenge and submitted a seven distinct recognition approaches that ranged from descriptor-based methods and deep-learning models to ensemble techniques that relied on multiple data representations to maximize performance and minimize bias. A comprehensive investigation into the performance of the submitted models is presented, as well as an in-depth analysis of bias and associated performance differentials due to differences in gender and ethnicity. The results of the challenge suggest that a wide variety of models (e.g., transformers, convolutional neural networks, ensemble models) is capable of achieving competitive recognition results, but also that all of the models still exhibit considerable performance differentials with respect to both gender and ethnicity. To promote further development of unbiased and effective ear recognition models, the starter kit of UERC 2023 together with the baseline model, and training and test data is made available from: http://ears.fri.uni-lj.si/ Ziga Emersic, Tetsushi Ohki, Muku Akasaka, Takahiko Arakawa, Soshi Maeda, Masora Okano, Yuya Sato, Anjith George, Sébastien Marcel, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Sajid Javed, Naoufel Werghi, S. G. Isik, Erdi Saritas, Hazim Kemal Ekenel, V. Hudovernik, Jan Niklas Kolf, Fadi Boutros, Naser Damer, G. Sharma, Aman Kamboj, Aditya Nigam, Deepak Kumar Jain 0001, G. Cámara-Chávez, Peter Peer, Vitomir Struc |
IJCB | 27 |
| 2023 | EFaR 2023: Efficient Face Recognition CompetitionabstractThis paper presents the summary of the Efficient Face Recognition Competition (EFaR) held at the 2023 International Joint Conference on Biometrics (IJCB 2023). The competition received 17 submissions from 6 different teams. To drive further development of efficient face recognition models, the submitted solutions are ranked based on a weighted score of the achieved verification accuracies on a diverse set of benchmarks, as well as the deployability given by the number of floating-point operations and model size. The evaluation of submissions is extended to bias, cross-quality, and large-scale recognition benchmarks. Overall, the paper gives an overview of the achieved performance values of the submitted solutions as well as a diverse set of baselines. The submitted solutions use small, efficient network architectures to reduce the computational cost, some solutions apply model quantization. An outlook on possible techniques that are underrepresented in current solutions is given as well. Jan Niklas Kolf, Fadi Boutros, Jurek Elliesen, Markus Theuerkauf, Naser Damer, Mohamad Alansari, Oussama Abdul Hay, Sara Alansari, Sajid Javed, Naoufel Werghi, Klemen Grm, Vitomir Struc, Fernando Alonso-Fernandez, Kevin Hernandez-Diaz, Josef Bigün, Anjith George, Christophe Ecabert, Hatef Otroshi-Shahreza, Ketan Kotwal, Sébastien Marcel, Iurii Medvedev, Bo Jin 0018, Diogo Nunes, Ahmad Hassanpour, Pankaj Khatiwada, Aafan Ahmad Toor, Bian Yang |
IJCB | 12 |
| 2023 | DFGC-VRA: DeepFake Game Competition on Visual Realism AssessmentabstractThis paper presents the summary report on the DeepFake Game Competition on Visual Realism Assessment (DFGC-VRA). Deep-learning based face-swap videos, also known as deepfakes, are becoming more and more realistic and deceiving. The malicious usage of these face-swap videos has caused wide concerns. There is a ongoing deepfake game between its creators and detectors, with the human in the loop. The research community has been focusing on the automatic detection of these fake videos, but the assessment of their visual realism, as perceived by human eyes, is still an unexplored dimension. Visual realism assessment, or VRA, is essential for assessing the potential impact that may be brought by a specific face-swap video, and it is also useful as a quality metric to compare different face-swap methods. This is the third edition of DFGC competitions, which focuses on the new visual realism assessment topic, different from previous ones that compete creators versus detectors. With this competition, we conduct a comprehensive study of the SOTA performance on the new task. We also release our MindSpore codes to further facilitate research in this field (https://github.com/bomb2peng/DFGC-VRA-benckmark). Bo Peng 0002, Xianyun Sun, Caiyong Wang, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Rongyu Zhang, Heng Cong, Lingzhi Fu, Yusheng Zhang, Boyuan Liu, Luka Dragar, Borut Batagelj, Peter Peer, Vitomir Struc, Xinghui Zhou, Kunlin Liu, Wenxiu Diao |
IJCB | 19 |
| 2023 | SeeABLE: Soft Discrepancies and Bounded Contrastive Learning for Exposing DeepfakesabstractModern deepfake detectors have achieved encouraging results, when training and test images are drawn from the same data collection. However, when these detectors are applied to images produced with unknown deepfake-generation techniques, considerable performance degradations are commonly observed. In this paper, we propose a novel deepfake detector, called SeeABLE, that formalizes the detection problem as a (one-class) out-of-distribution detection task and generalizes better to unseen deepfakes. Specifically, SeeABLE first generates local image perturbations (referred to as soft-discrepancies) and then pushes the perturbed faces towards predefined prototypes using a novel regression-based bounded contrastive loss. To strengthen the generalization performance of SeeABLE to unknown deepfake types, we generate a rich set of soft discrepancies and train the detector: (i) to localize, which part of the face was modified, and (ii) to identify the alteration type. To demonstrate the capabilities of SeeABLE, we perform rigorous experiments on several widely-used deepfake datasets and show that our model convincingly outperforms competing state-of-the-art detectors, while exhibiting highly encouraging generalization capabilities. The source code for SeeABLE is available from: https://github.com/anonymous-author-sub/seeable. Nicolas Larue, Ngoc-Son Vu, Vitomir Struc, Peter Peer, Vassilis Christophides |
ICCV | 3 |
| 2023 | Synthetic data for face recognition: Current state and future prospects
Fadi Boutros, Vitomir Struc, Julian Fierrez, Naser Damer |
Image Vis. Comput. | 2 |
| 2023 | A survey on computer vision based human analysis in the COVID-19 era
Fevziye Irem Eyiokur, Alperen Kantarci, Mustafa Ekrem Erakin, Naser Damer, Ferda Ofli, Muhammad Imran 0002, Janez Krizaj, Albert Ali Salah, Alex Waibel, Vitomir Struc, Hazim Kemal Ekenel |
Image Vis. Comput. | 10 |
| 2023 | Face deidentification with controllable privacy protectionabstractPrivacy protection has become a crucial concern in today’s digital age. Particularly sensitive here are facial images, which typically not only reveal a person’s identity, but also other sensitive personal information. To address this problem, various face deidentification techniques have been presented in the literature. These techniques try to remove or obscure personal information from facial images while still preserving their usefulness for further analysis. While a considerable amount of work has been proposed on face deidentification, most state-of-the-art solutions still suffer from various drawbacks, and (a) deidentify only a narrow facial area, leaving potentially important contextual information unprotected, (b) modify facial images to such degrees, that image naturalness and facial diversity is suffering in the deidentify images, (c) offer no flexibility in the level of privacy protection ensured, leading to suboptimal deployment in various applications, and (d) often offer an unsatisfactory trade-off between the ability to obscure identity information, quality and naturalness of the deidentified images, and sufficient utility preservation. In this paper, we address these shortcomings with a novel controllable face deidentification technique that balances image quality, identity protection, and data utility for further analysis. The proposed approach utilizes a powerful generative model (StyleGAN2), multiple auxiliary classification models, and carefully designed constraints to guide the deidentification process. The approach is validated across four diverse datasets (CelebA-HQ, RaFD, XM2VTS, AffectNet) and in comparison to 7 state-of-the-art competitors. The results of the experiments demonstrate that the proposed solution leads to: (a) a considerable level of identity protection, (b) valuable preservation of data utility, (c) sufficient diversity among the deidentified faces, and (d) encouraging overall performance. Blaz Meden, Manfred Gonzalez-Hernandez, Peter Peer, Vitomir Struc |
Image Vis. Comput. | 4 |
| 2023 | Exploring Bias in Sclera Segmentation Models: A Group Evaluation ApproachabstractBias and fairness of biometric algorithms have been key topics of research in recent years, mainly due to the societal, legal and ethical implications of potentially unfair decisions made by automated decision-making models. A considerable amount of work has been done on this topic across different biometric modalities, aiming at better understanding the main sources of algorithmic bias or devising mitigation measures. In this work, we contribute to these efforts and present the first study investigating bias and fairness of sclera segmentation models. Although sclera segmentation techniques represent a key component of sclera-based biometric systems with a considerable impact on the overall recognition performance, the presence of different types of biases in sclera segmentation methods is still underexplored. To address this limitation, we describe the results of a group evaluation effort (involving seven research groups), organized to explore the performance of recent sclera segmentation models within a common experimental framework and study performance differences (and bias), originating from various demographic as well as environmental factors. Using five diverse datasets, we analyze seven independently developed sclera segmentation models in different experimental configurations. The results of our experiments suggest that there are significant differences in the overall segmentation performance across the seven models and that among the considered factors, ethnicity appears to be the biggest cause of bias. Additionally, we observe that training with representative and balanced data does not necessarily lead to less biased results. Finally, we find that in general there appears to be a negative correlation between the amount of bias observed (due to eye color, ethnicity and acquisition device) and the overall segmentation performance, suggesting that advances in the field of semantic segmentation may also help with mitigating bias. Matej Vitek, Abhijit Das 0001, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Jalil Nourmohammadi-Khiarak, Mohsen Akbari Shahpar, Meysam Asgari-Chenaghlu, Farhang Jaryani, Juan E. Tapia, Andres Valenzuela, Caiyong Wang, Yunlong Wang 0003, Zhaofeng He 0001, Zhenan Sun, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Kiran B. Raja, Gourav Gupta, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, S. V. Aruna Kumar, B. S. Harish, Umapada Pal 0001, Peter Peer, Vitomir Struc |
IEEE Trans. Inf. Forensics Secur. | 29 |
| 2023 | MaskFaceGAN: High-Resolution Face Editing With Masked GAN Latent Code OptimizationabstractFace editing represents a popular research topic within the computer vision and image processing communities. While significant progress has been made recently in this area, existing solutions: (i) are still largely focused on low-resolution images, (ii) often generate editing results with visual artefacts, or (iii) lack fine-grained control over the editing procedure and alter multiple (entangled) attributes simultaneously, when trying to generate the desired facial semantics. In this paper, we aim to address these issues through a novel editing approach, called MaskFaceGAN that focuses on local attribute editing. The proposed approach is based on an optimization procedure that directly optimizes the latent code of a pre-trained (state-of-the-art) Generative Adversarial Network (i.e., StyleGAN2) with respect to several constraints that ensure: (i) preservation of relevant image content, (ii) generation of the targeted facial attributes, and (iii) spatially-selective treatment of local image regions. The constraints are enforced with the help of an (differentiable) attribute classifier and face parser that provide the necessary reference information for the optimization procedure. MaskFaceGAN is evaluated in extensive experiments on the FRGC, SiblingsDB-HQf, and XM2VTS datasets and in comparison with several state-of-the-art techniques from the literature. Our experimental results show that the proposed approach is able to edit face images with respect to several local facial attributes with unprecedented image quality and at high-resolutions ( 1024×1024 ), while exhibiting considerably less problems with attribute entanglement than competing solutions. The source code is publicly available from: https://github.com/MartinPernus/MaskFaceGAN. Martin Pernus, Vitomir Struc, Simon Dobrisek |
IEEE Trans. Image Process. | 2 |
| 2022 | SYN-MAD 2022: Competition on Face Morphing Attack Detection Based on Privacy-aware Synthetic Training DataabstractThis paper presents a summary of the Competition on Face Morphing Attack Detection Based on Privacy-aware Synthetic Training Data (SYN-MAD) held at the 2022 In-ternational Joint Conference on Biometrics (IJCB 2022). The competition attracted a total of 12 participating teams, both from academia and industry and present in 11 differ-ent countries. In the end, seven valid submissions were submitted by the participating teams and evaluated by the organizers. The competition was held to present and at-tract solutions that deal with detecting face morphing at-tacks while protecting people's privacy for ethical and le-gal reasons. To ensure this, the training data was limited to synthetic data provided by the organizers. The submitted solutions presented innovations that led to out-performing the considered baseline in many experimental settings. The evaluation benchmark is now available at: https://github.com/marcohuber/SYN-MAD-2022. Marco Huber, Fadi Boutros, Anh Thi Luu, Kiran B. Raja, Ramachandra Raghavendra, Naser Damer, Pedro C. Neto, Tiago Gonçalves 0001, Ana Filipa Sequeira, Jaime S. Cardoso 0001, João Tremoço, Miguel Lourenço, Sergio Serra, Eduardo Cermeño, Marija Ivanovska, Borut Batagelj, Andrej Kronovsek, Peter Peer, Vitomir Struc |
IJCB | 19 |
| 2022 | BiOcularGAN: Bimodal Synthesis and Annotation of Ocular ImagesabstractCurrent state-of-the-art segmentation techniques for ocular images are critically dependent on large-scale annotated datasets, which are labor-intensive to gather and often raise privacy concerns. In this paper, we present a novel framework, called BiOcularGAN, capable of generating synthetic large-scale datasets of photorealistic (visible light and near-infrared) ocular images, together with corresponding segmentation labels to address these issues. At its core, the framework relies on a novel Dual-Branch StyleGAN2 (DB-StyleGAN2) model that facilitates bimodal image generation, and a Semantic Mask Generator (SMG) component that produces semantic annotations by exploiting latent features of the DB-StyleGAN2 model. We evaluate BiOcularGAN through extensive experiments across five diverse ocular datasets and analyze the effects of bi-modal data generation on image quality and the produced annotations. Our experimental results show that BiOcularGAN is able to produce high-quality matching bimodal images and annotations (with minimal manual intervention) that can be used to train highly competitive (deep) segmentation models (in a privacy aware-manner) that perform well across multiple real-world datasets. The source code for the BiOcularGAN framework is publicly available at https://github.com/dariant/BiOcularGAN. Darian Tomasevic, Peter Peer, Vitomir Struc |
IJCB | 3 |
| 2022 | FaceQAN: Face Image Quality Assessment Through Adversarial Noise ExplorationabstractRecent state-of-the-art face recognition (FR) approaches have achieved impressive performance, yet unconstrained face recognition still represents an open problem. Face image quality assessment (FIQA) approaches aim to estimate the quality of the input samples that can help provide information on the confidence of the recognition decision and eventually lead to improved results in challenging scenarios. While much progress has been made in face image quality assessment in recent years, computing reliable quality scores for diverse facial images and FR models remains challenging. In this paper, we propose a novel approach to face image quality assessment, called FaceQAN, that is based on adversarial examples and relies on the analysis of adversarial noise which can be calculated with any FR model learned by using some form of gradient descent. As such, the proposed approach is the first to link image quality to adversarial attacks. Comprehensive (cross-model as well as model-specific) experiments are conducted with four benchmark datasets, i.e., LFW, CFP–FP, XQLFW and IJB–C, four FR models, i.e., CosFace, ArcFace, CurricularFace and ElasticFace, and in comparison to seven state-of-the-art FIQA methods to demonstrate the performance of FaceQAN. Experimental results show that FaceQAN achieves competitive results, while exhibiting several desirable characteristics. The source code for FaceQAN is available at https://github.com/LSIbabnikz/FaceQAN. Ziga Babnik, Peter Peer, Vitomir Struc |
ICPR | 3 |
| 2022 | C-VTON: Context-Driven Image-Based Virtual Try-On NetworkabstractImage-based virtual try-on techniques have shown great promise for enhancing the user-experience and improving customer satisfaction on fashion-oriented e-commerce platforms. However, existing techniques are currently still limited in the quality of the try-on results they are able to produce from input images of diverse characteristics. In this work, we propose a Context-Driven Virtual Try-On Network (C-VTON) that addresses these limitations and convincingly transfers selected clothing items to the target subjects even under challenging pose configurations and in the presence of self-occlusions. At the core of the C-VTON pipeline are: (i) a geometric matching procedure that efficiently aligns the target clothing with the pose of the person in the input images, and (ii) a powerful image generator that utilizes various types of contextual information when synthesizing the final try-on result. C-VTON is evaluated in rigorous experiments on the VITON and MPV datasets and in comparison to state-of-the-art techniques from the literature. Experimental results show that the proposed approach is able to produce photo-realistic and visually convincing results and significantly improves on the existing state-of-the-art. Benjamin Fele, Ajda Lampe, Peter Peer, Vitomir Struc |
WACV | 4 |
| 2022 | DHF-Net: A hierarchical feature interactive fusion network for dialogue emotion recognitionabstractTo balance the trade-off between contextual information and fine-grained information in identifying specific emotions during a dialogue and combine the interaction of hierarchical feature related information, this paper proposes a hierarchical feature interactive fusion network (named DHF-Net), which not only can retain the integrity of the context sequence information but also can extract more fine-grained information. To obtain a deep semantic information, DHF-Net processes the task of recognizing dialogue emotion and dialogue act/intent separately, and then learns the cross-impact of two tasks through collaborative attention. Also, a bidirectional gate recurrent unit (Bi-GRU) connected hybrid convolutional neural network (CNN) group method is designed, by which the sequence information is smoothly sent to the multi-level local information layers for feature exaction. Experimental results show that, on two open session datasets, the performance of DHF-Net is improved by 1.8% and 1.2%, respectively. Chenquan Gan, Yucheng Yang 0007, Qingyi Zhu, Deepak Kumar Jain 0001, Vitomir Struc |
Expert Syst. Appl. | 5 |
| 2021 | MFR 2021: Masked Face Recognition CompetitionabstractThis paper presents a summary of the Masked Face Recognition Competitions (MFR) held within the 2021 International Joint Conference on Biometrics (IJCB 2021). The competition attracted a total of 10 participating teams with valid submissions. The affiliations of these teams are diverse and associated with academia and industry in nine different countries. These teams successfully submitted 18 valid solutions. The competition is designed to motivate solutions aiming at enhancing the face recognition accuracy of masked faces. Moreover, the competition considered the deployability of the proposed solutions by taking the compactness of the face recognition models into account. A private dataset representing a collaborative, multisession, real masked, capture scenario is used to evaluate the submitted solutions. In comparison to one of the topperforming academic face recognition solutions, 10 out of the 18 submitted solutions did score higher masked face verification accuracy. Fadi Boutros, Naser Damer, Jan Niklas Kolf, Kiran B. Raja, Florian Kirchbuchner, Ramachandra Raghavendra, Arjan Kuijper, Pengcheng Fang, Fei Wang 0032, David Montero 0002, Naiara Aginako, Basilio Sierra, Marcos Nieto Doncel, Mustafa Ekrem Erakin, Ugur Demir, Hazim Kemal Ekenel, Asaki Kataoka, Kohei Ichikawa, Shizuma Kubo, Jie Zhang 0071, Shiguang Shan, Klemen Grm, Vitomir Struc, Sachith Seneviratne, Nuran Kasthuriarachchi, Sanka Rasnayaka, Pedro C. Neto, Ana Filipa Sequeira, João Ribeiro Pinto, Mohsen Saffari, Jaime S. Cardoso 0001 |
IJCB | 26 |
| 2021 | NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and LocalizationabstractFor iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research. Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad |
IJCB | 56 |
| 2021 | Compact Finite-State Super Transducers for Grapheme-to-Phoneme Conversion in Highly Inflected Languages
Ziga Golob, Bostjan Vesnicer, Mario Zganec, Vitomir Struc, Simon Dobrisek, Jerneja Zganec-Gros |
ICIC (1) | 4 |
| 2021 | Privacy-Enhancing Face Biometrics: A Comprehensive SurveyabstractBiometric recognition technology has made significant advances over the last decade and is now used across a number of services and applications. However, this widespread deployment has also resulted in privacy concerns and evolving societal expectations about the appropriate use of the technology. For example, the ability to automatically extract age, gender, race, and health cues from biometric data has heightened concerns about privacy leakage. Face recognition technology, in particular, has been in the spotlight, and is now seen by many as posing a considerable risk to personal privacy. In response to these and similar concerns, researchers have intensified efforts towards developing techniques and computational models capable of ensuring privacy to individuals, while still facilitating the utility of face recognition technology in several application scenarios. These efforts have resulted in a multitude of privacy-enhancing techniques that aim at addressing privacy risks originating from biometric systems and providing technological solutions for legislative requirements set forth in privacy laws and regulations, such as GDPR. The goal of this overview paper is to provide a comprehensive introduction into privacy-related research in the area of biometrics and review existing work on Biometric Privacy-Enhancing Techniques (B-PETs) applied to face biometrics. To make this work useful for as wide of an audience as possible, several key topics are covered as well, including evaluation strategies used with B-PETs, existing datasets, relevant standards, and regulations and critical open issues that will have to be addressed in the future. Blaz Meden, Peter Rot, Philipp Terhörst, Naser Damer, Arjan Kuijper, Walter J. Scheirer, Arun Ross, Peter Peer, Vitomir Struc |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2020 | Learning privacy-enhancing face representations through feature disentanglementabstractConvolutional Neural Networks (CNNs) are today the de-facto standard for extracting compact and discriminative face representations (templates) from images in automatic face recognition systems. Due to the characteristics of CNN models, the generated representations typically encode a multitude of information ranging from identity to soft-biometric attributes, such as age, gender or ethnicity. However, since these representations were computed for the purpose of identity recognition only, the soft-biometric information contained in the templates represents a serious privacy risk. To mitigate this problem, we present in this paper a privacy-enhancing approach capable of suppressing potentially sensitive soft-biometric information in face representations without significantly compromising identity information. Specifically, we introduce a Privacy-Enhancing Face-Representation learning Network (PFRNet) that disentangles identity from attribute information in face representations and consequently allows to efficiently suppress soft-biometrics in face templates. We demonstrate the feasibility of PFRNet on the problem of gender suppression and show through rigorous experiments on the CelebA, Labeled Faces in the Wild (LFW) and Adience datasets that the proposed disentanglement-based approach is highly effective and improves significantly on the existing state-of-the-art. Blaz Bortolato, Marija Ivanovska, Peter Rot, Janez Krizaj, Philipp Terhörst, Naser Damer, Peter Peer, Vitomir Struc |
FG | 8 |
| 2020 | SSBC 2020: Sclera Segmentation Benchmarking Competition in the Mobile EnvironmentabstractThe paper presents a summary of the 2020 Sclera Segmentation Benchmarking Competition (SSBC), the 7th in the series of group benchmarking efforts centred around the problem of sclera segmentation. Different from previous editions, the goal of SSBC 2020 was to evaluate the performance of sclera-segmentation models on images captured with mobile devices. The competition was used as a platform to assess the sensitivity of existing models to i) differences in mobile devices used for image capture and ii) changes in the ambient acquisition conditions. 26 research groups registered for SSBC 2020, out of which 13 took part in the final round and submitted a total of 16 segmentation models for scoring. These included a wide variety of deep-learning solutions as well as one approach based on standard image processing techniques. Experiments were conducted with three recent datasets. Most of the segmentation models achieved relatively consistent performance across images captured with different mobile devices (with slight differences across devices), but struggled most with low-quality images captured in challenging ambient conditions, i.e., in an indoor environment and with poor lighting. Matej Vitek, Abhijit Das 0001, Yann Pourcenoux, Alexandre Missler, C. Paumier, Sumanta Das, Ishita De Ghosh, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Junxing Hu, Yong He 0009, Caiyong Wang, Yunlong Wang 0003, Zhenan Sun, Dailé Osorio Roig, Christian Rathgeb, Christoph Busch 0001, Juan E. Tapia, Andres Valenzuela, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, Sabari Nathan, R. Suganya 0001, Vineet Mehta, Abhinav Dhall, Kiran B. Raja, Gourav Gupta, Jalil Nourmohammadi-Khiarak, Mohsen Akbari-Shahper, Farhang Jaryani, Meysam Asgari-Chenaghlu, Ritesh Vyas, Sristi Dakshit, Peter Peer, Umapada Pal 0001, Vitomir Struc |
IJCB | 43 |
| 2020 | Evaluation and analysis of ear recognition models: performance, complexity and resource requirements
Ziga Emersic, Blaz Meden, Peter Peer, Vitomir Struc |
Neural Comput. Appl. | 4 |
| 2020 | Simultaneous multi-descent regression and feature learning for facial landmarking in depth imagesabstractAbstract Face alignment (or facial landmarking) is an important task in many face-related applications, ranging from registration, tracking, and animation to higher-level classification problems such as face, expression, or attribute recognition. While several solutions have been presented in the literature for this task so far, reliably locating salient facial features across a wide range of posses still remains challenging. To address this issue, we propose in this paper a novel method for automatic facial landmark localization in 3D face data designed specifically to address appearance variability caused by significant pose variations. Our method builds on recent cascaded regression-based methods to facial landmarking and uses a gating mechanism to incorporate multiple linear cascaded regression models each trained for a limited range of poses into a single powerful landmarking model capable of processing arbitrary-posed input data. We develop two distinct approaches around the proposed gating mechanism: (1) the first uses a gated multiple ridge descent mechanism in conjunction with established (hand-crafted) histogram of gradients features for face alignment and achieves state-of-the-art landmarking performance across a wide range of facial poses and (2) the second simultaneously learns multiple-descent directions as well as binary features that are optimal for the alignment tasks and in addition to competitive landmarking results also ensures extremely rapid processing. We evaluate both approaches in rigorous experiments on several popular datasets of 3D face images, i.e., the FRGCv2 and Bosphorus 3D face datasets and image collections F and G from the University of Notre Dame. The results of our evaluation show that both approaches compare favorably to the state-of-the-art, while exhibiting considerable robustness to pose variations. Janez Krizaj, Peter Peer, Vitomir Struc, Simon Dobrisek |
Neural Comput. Appl. | 3 |
| 2020 | A comprehensive investigation into sclera biometrics: a novel dataset and performance study
Matej Vitek, Peter Rot, Vitomir Struc, Peter Peer |
Neural Comput. Appl. | 3 |
| 2020 | Face Hallucination Using Cascaded Super-Resolution and Identity PriorsabstractIn this paper we address the problem of hallucinating high-resolution facial images from low-resolution inputs at high magnification factors. We approach this task with convolutional neural networks (CNNs) and propose a novel (deep) face hallucination model that incorporates identity priors into the learning procedure. The model consists of two main parts: i) a cascaded super-resolution network that upscales the low-resolution facial images, and ii) an ensemble of face recognition models that act as identity priors for the super-resolution network during training. Different from most competing super-resolution techniques that rely on a single model for upscaling (even with large magnification factors), our network uses a cascade of multiple SR models that progressively upscale the low-resolution images using steps of 2× . This characteristic allows us to apply supervision signals (target appearances) at different resolutions and incorporate identity constraints at multiple-scales. The proposed C-SRIP model (Cascaded Super Resolution with Identity Priors) is able to upscale (tiny) low-resolution images captured in unconstrained conditions and produce visually convincing results for diverse low-resolution inputs. We rigorously evaluate the proposed model on the Labeled Faces in the Wild (LFW), Helen and CelebA datasets and report superior performance compared to the existing state-of-the-art. Klemen Grm, Walter J. Scheirer, Vitomir Struc |
IEEE Trans. Image Process. | 3 |
| 2019 | Frame-based classification for cross-speed gait recognition
Jure Kovac, Vitomir Struc, Peter Peer |
Multim. Tools Appl. | 2 |
| 2018 | To Frontalize or Not to Frontalize: Do We Really Need Elaborate Pre-processing to Improve Face Recognition?abstractFace recognition performance has improved remarkably in the last decade. Much of this success can be attributed to the development of deep learning techniques such as convolutional neural networks (CNNs). While CNNs have pushed the state-of-the-art forward, their training process requires a large amount of clean and correctly labelled training data. If a CNN is intended to tolerate facial pose, then we face an important question: should this training data be diverse in its pose distribution, or should face images be normalized to a single pose in a pre-processing step? To address this question, we evaluate a number of facial landmarking algorithms and a popular frontalization method to understand their effect on facial recognition performance. Additionally, we introduce a new, automatic, single-image frontalization scheme that exceeds the performance of the reference frontalization algorithm for video-to-video face matching on the Point and Shoot Challenge (PaSC) dataset. Additionally, we investigate failure modes of each frontalization method on different facial yaw using the CMU Multi-PIE dataset. We assert that the subsequent recognition and verification performance serves to quantify the effectiveness of each pose correction scheme. Sandipan Banerjee, Joel Brogan, Janez Krizaj, Aparna Bharati, Brandon RichardWebster, Vitomir Struc, Patrick J. Flynn, Walter J. Scheirer |
WACV | 6 |
| 2018 | UG^2: A Video Benchmark for Assessing the Impact of Image Restoration and Enhancement on Automatic Visual RecognitionabstractAdvances in image restoration and enhancement techniques have led to discussion about how such algorithms can be applied as a pre-processing step to improve automatic visual recognition. In principle, techniques like deblurring and super-resolution should yield improvements by de-emphasizing noise and increasing signal in an input image. But the historically divergent goals of computational photography and visual recognition communities have created a significant need for more work in this direction. To facilitate new research, we introduce a new benchmark dataset called UG2, which contains three difficult real-world scenarios: uncontrolled videos taken by UAVs and manned gliders, as well as controlled videos taken on the ground. Over 150,000 annotated frames for hundreds of ImageNet classes are available, which are used for baseline experiments that assess the impact of known and unknown image artifacts and other conditions on common deep learning-based object classification approaches. Further, current image restoration and enhancement techniques are evaluated by determining whether or not they improve baseline classification performance. Results show that there is plenty of room for algorithmic innovation, making this dataset a useful tool going forward. Rosaura G. VidalMata, Sreya Banerjee, Klemen Grm, Vitomir Struc, Walter J. Scheirer |
WACV | 4 |
| 2017 | Foreword
Bir Bhanu, Abdenour Hadid, Mark Nixon, Vitomir Struc |
FG | 5 |
| 2017 | Training Convolutional Neural Networks with Limited Training Data for Ear Recognition in the WildabstractIdentity recognition from ear images is an active field of research within the biometric community. The ability to capture ear images from a distance and in a covert manner makes ear recognition technology an appealing choice for surveillance and security applications as well as related application domains. In contrast to other biometric modalities, where large datasets captured in uncontrolled settings are readily available, datasets of ear images are still limited in size and mostly of laboratory-like quality. As a consequence, ear recognition technology has not benefited yet from advances in deep learning and convolutionalneural networks (CNNs) and is still lacking behind other modalities that experienced significant performance gains owing to deep recognition technology. In this paper we address this problem and aim at building a CNN-based ear recognition model. We explore different strategies towards model training with limited amounts of training data and show that by selecting an appropriate model architecture, using aggressive data augmentation and selective learning on existing (pre-trained) models, we are able to learn an effective CNN·based model using a little more than 1300training images. The result of our work is the first CNN·based approach to ear recognition that is also made publicly available to the research community. With our model we are able to improve on the rank one recognition rate of the previous state-of-the-art by more than 25% on a challenging dataset of ear images captured from the web (a.k.a, in the wild). Ziga Emersic, Dejan Stepec, Vitomir Struc, Peter Peer |
FG | 3 |
| 2017 | SSERBC 2017: Sclera segmentation and eye recognition benchmarking competitionabstractThis paper summarises the results of the Sclera Segmentation and Eye Recognition Benchmarking Competition (SSERBC 2017). It was organised in the context of the International Joint Conference on Biometrics (IJCB 2017). The aim of this competition was to record the recent developments in sclera segmentation and eye recognition in the visible spectrum (using iris, sclera and peri-ocular, and their fusion), and also to gain the attention of researchers on this subject. In this regard, we have used the Multi-Angle Sclera Dataset (MASD version 1). It is comprised of2624 images taken from both the eyes of 82 identities. Therefore, it consists of images of 164 (82×2) eyes. A manual segmentation mask of these images was created to baseline both tasks. Precision and recall based statistical measures were employed to evaluate the effectiveness of the segmentation and the ranks of the segmentation task. Recognition accuracy measure has been employed to measure the recognition task. Manually segmented sclera, iris and peri-ocular regions were used in the recognition task. Sixteen teams registered for the competition, and among them, six teams submitted their algorithms or systems for the segmentation task and two of them submitted their recognition algorithm or systems. The results produced by these algorithms or systems reflect current developments in the literature of sclera segmentation and eye recognition, employing cutting edge techniques. The MASD version 1 dataset with some of the ground truth will be freely available for research purposes. The success of the competition also demonstrates the recent interests of researchers from academia as well as industry on this subject. Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein, Dejan Stepec, Peter Rot, Ziga Emersic, Peter Peer, Vitomir Struc, S. V. Aruna Kumar, B. S. Harish |
IJCB | 9 |
| 2017 | The unconstrained ear recognition challengeabstractIn this paper we present the results of the Unconstrained Ear Recognition Challenge (UERC), a group benchmarking effort centered around the problem of person recognition from ear images captured in uncontrolled conditions. The goal of the challenge was to assess the performance of existing ear recognition techniques on a challenging large-scale dataset and identify open problems that need to be addressed in the future. Five groups from three continents participated in the challenge and contributed six ear recognition techniques for the evaluation, while multiple baselines were made available for the challenge by the UERC organizers. A comprehensive analysis was conducted with all participating approaches addressing essential research questions pertaining to the sensitivity of the technology to head rotation, flipping, gallery size, large-scale recognition and others. The top performer of the UERC was found to ensure robust performance on a smaller part of the dataset (with 180 subjects) regardless of image characteristics, but still exhibited a significant performance drop when the entire dataset comprising 3,704 subjects was used for testing. Ziga Emersic, Dejan Stepec, Vitomir Struc, Peter Peer, Anjith George, Adil M. Ahmad, Elshibani Omar, Terrance E. Boult, Reza Safdari, Stefanos Zafeiriou, Doggucan Yaman, Fevziye Irem Eyiokur, Hazim Kemal Ekenel |
IJCB | 3 |
| 2017 | Face deidentification with generative deep neural networksabstractFace deidentification is an active topic amongst privacy and security researchers. Early deidentification methods relying on image blurring or pixelisation have been replaced in recent years with techniques based on formal anonymity models that provide privacy guaranties and retain certain characteristics of the data even after deidentification. The latter aspect is important, as it allows the deidentified data to be used in applications for which identity information is irrelevant. In this work, the authors present a novel face deidentification pipeline, which ensures anonymity by synthesising artificial surrogate faces using generative neural networks (GNNs). The generated faces are used to deidentify subjects in images or videos, while preserving non‐identity‐related aspects of the data and consequently enabling data utilisation. Since generative networks are highly adaptive and can utilise diverse parameters (pertaining to the appearance of the generated output in terms of facial expressions, gender, race etc.), they represent a natural choice for the problem of face deidentification. To demonstrate the feasibility of the authors’ approach, they perform experiments using automated recognition tools and human annotators. Their results show that the recognition performance on deidentified images is close to chance, suggesting that the deidentification process based on GNNs is effective. Blaz Meden, Refik Can Malli, Sebastjan Fabijan, Hazim Kemal Ekenel, Vitomir Struc, Peter Peer |
IET Signal Process. | 5 |
| 2017 | Ear recognition: More than a survey
Ziga Emersic, Vitomir Struc, Peter Peer |
Neurocomputing | 2 |
| 2014 | The IJCB 2014 PaSC video face and person recognition competitionabstractThe Point-and-Shoot Face Recognition Challenge (PaSC) is a performance evaluation challenge including 1401 videos of 265 people acquired with handheld cameras and depicting people engaged in activities with non-frontal head pose. This report summarizes the results from a competition using this challenge problem. In the Video-to-video Experiment a person in a query video is recognized by comparing the query video to a set of target videos. Both target and query videos are drawn from the same pool of 1401 videos. In the Still-to-video Experiment the person in a query video is to be recognized by comparing the query video to a larger target set consisting of still images. Algorithm performance is characterized by verification rate at a false accept rate of 0.01 and associated receiver operating characteristic (ROC) curves. Participants were provided eye coordinates for video frames. Results were submitted by 4 institutions: (i) Advanced Digital Science Center, Singapore; (ii) CPqD, Brasil; (iii) Stevens Institute of Technology, USA; and (iv) University of Ljubljana, Slovenia. Most competitors demonstrated video face recognition performance superior to the baseline provided with PaSC. The results represent the best performance to date on the handheld video portion of the PaSC. J. Ross Beveridge, Hao Zhang 0013, Patrick J. Flynn, Yooyoung Lee, Venice Erin Liong, Jiwen Lu, Marcus A. Angeloni, Tiago de Freitas Pereira, Gang Hua 0001, Vitomir Struc, Janez Krizaj, P. Jonathon Phillips |
IJCB | 11 |
| 2013 | Patch-wise low-dimensional probabilistic linear discriminant analysis for Face RecognitionabstractThe paper introduces a novel approach to face recognition based on the recently proposed low-dimensional probabilistic linear discriminant analysis (LD-PLDA). The proposed approach is specifically designed for complex recognition tasks, where highly nonlinear face variations are typically encountered. Such data variations are commonly induced by changes in the external illumination conditions, viewpoint changes or expression variations and represent quite a challenge even for state-of-the-art techniques, such as LD-PLDA. To overcome this problem, we propose here a patch-wise form of the LDPLDA technique (i.e., PLD-PLDA), which relies on local image patches rather than the entire image to make inferences about the identity of the input images. The basic idea here is to decompose the complex face recognition problem into simpler problems, for which the linear nature of the LD-PLDA technique may be better suited. By doing so, several similarity scores are derived from one facial image, which are combined at the final stage using a simple sum-rule fusion scheme to arrive at a single score that can be employed for identity inference. We evaluate the proposed technique on experiment 4 of the Face Recognition Grand Challenge (FRGCv2) database with highly promising results. Vitomir Struc, Nikola Pavesic, Jerneja Zganec-Gros, Bostjan Vesnicer |
ICASSP | 1 |
| 2012 | Non-parametric score normalization for biometric verification systems
Vitomir Struc, Jerneja Zganec-Gros, Nikola Pavesic |
ICPR | 1 |
| 2010 | Removing illumination artifacts from face images using the nuisance attribute projectionabstractIllumination induced appearance changes represent one of the open challenges in automated face recognition systems still significantly influencing their performance. Several techniques have been presented in the literature to cope with this problem; however, a universal solution remains to be found. In this paper we present a novel normalization scheme based on the nuisance attribute projection (NAP), which tries to remove the effects of illumination by projecting away multiple dimensions of a low dimensional illumination subspace. The technique is assessed in face recognition experiments performed on the extended YaleB and XM2VTS databases. Comparative results with state-of-the-art techniques show the competitiveness of the proposed technique. Vitomir Struc, Bostjan Vesnicer, France Mihelic, Nikola Pavesic |
ICASSP | 1 |
| 2010 | Multi-modal Emotion Recognition Using Canonical Correlations and Acoustic FeaturesabstractThe information of the psycho-physical state of the subject is becoming a valuable addition to the modern audio or video recognition systems. As well as enabling a better user experience, it can also assist in superior recognition accuracy of the base system. In the article, we present our approach to multi-modal (audio-video) emotion recognition system. For audio sub-system, a feature set comprised of prosodic, spectral and cepstrum features is selected and support vector classifier is used to produce the scores for each emotional category. For video sub-system a novel approach is presented, which does not rely on the tracking of specific facial landmarks and thus, eliminates the problems usually caused, if the tracking algorithm fails at detecting the correct area. The system is evaluated on the eNTERFACE database and the recognition accuracy of our audio-video fusion is compared to the published results in the literature. Rok Gajsek, Vitomir Struc, France Mihelic |
ICPR | 2 |
| 2010 | Confidence Weighted Subspace Projection Techniques for Robust Face Recognition in the Presence of Partial OcclusionsabstractSubspace projection techniques are known to be susceptible to the presence of partial occlusions in the image data. To overcome this susceptibility, we present in this paper a confidence weighting scheme that assigns weights to pixels according to a measure, which quantifies the confidence that the pixel in question represents an outlier. With this procedure the impact of the occluded pixels on the subspace representation is reduced and robustness to partial occlusions is obtained. Next, the confidence weighting concept is improved by a local procedure for the estimation of the subspace representation. Both the global weighting approach and the local estimation procedure are assessed in face recognition experiments on the AR database, where encouraging results are obtained with partially occluded facial images. Vitomir Struc, Simon Dobrisek, Nikola Pavesic |
ICPR | 1 |
| 2010 | Gender and affect recognition based on GMM and GMM-UBM modeling with relevance MAP estimation
Rok Gajsek, Janez Zibert, Tadej Justin, Vitomir Struc, Bostjan Vesnicer, France Mihelic |
INTERSPEECH | 4 |
| 2010 | An Evaluation of Video-to-Video Face VerificationabstractPerson recognition using facial features, e.g., mug-shot images, has long been used in identity documents. However, due to the widespread use of web-cams and mobile devices embedded with a camera, it is now possible to realize facial video recognition, rather than resorting to just still images. In fact, facial video recognition offers many advantages over still image recognition; these include the potential of boosting the system accuracy and deterring spoof attacks. This paper presents an evaluation of person identity verification using facial video data, organized in conjunction with the International Conference on Biometrics (ICB 2009). It involves 18 systems submitted by seven academic institutes. These systems provide for a diverse set of assumptions, including feature representation and preprocessing variations, allowing us to assess the effect of adverse conditions, usage of quality information, query selection, and template construction for video-to-video face authentication. Norman Poh, Chi-Ho Chan, Josef Kittler, Sébastien Marcel, Chris McCool, Enrique Argones-Rúa, José Luis Alba-Castro, Mauricio Villegas, Roberto Paredes, Vitomir Struc, Nikola Pavesic, Albert Ali Salah, Hui Fang 0003, Nicholas Costen |
IEEE Trans. Inf. Forensics Secur. | 10 |
| 2009 | Emotion recognition using linear transformations in combination with videoabstractThe paper discuses the usage of linear transformations of Hidden Markov Models, normally employed for speaker and environment adaptation, as a way of extracting the emotional components from the speech. A constrained version of Maximum Likelihood Linear Regression (CMLLR) transformation is used as a feature for classification of normal or aroused emotional state. We present a procedure of incrementally building a set of speaker independent acoustic models, that are used to estimate the CMLLR transformations for emotion classification. An audio-video database of spontaneous emotions (AvID) is briefly presented since it forms the basis for the evaluation of the proposed method. Emotion classification using the video part of the database is also described and the added value of combining the visual information with the audio features is shown. Index Terms: emotion recognition, emotional database, linear transformations Rok Gajsek, Vitomir Struc, Simon Dobrisek, France Mihelic |
INTERSPEECH | 2 |